Skip to main content
Back to News
news/AI Infrastructure

Cloudflare Auto Router Cuts AI Inference Spend by Up to 30%

Cloudflare's new Auto Router, now in public beta, routes each AI request to the most cost-efficient capable model, cutting inference spend by up to 30 percent.

Stefan Trbojevic

Stefan Trbojevic

30 September 20262 min read
LinkedIn
Abstract illustration of an AI model routing hub with branching data pathways

The takeaway

Model routing is consolidating into a standard infrastructure layer. Cloudflare's Auto Router turns AI Gateway into a cost-control plane that picks the right model per task, not per token price.

Why it matters for builders

For teams shipping agents and internal AI tools, Auto Router removes the manual model-selection step from harnesses like OpenCode, Claude Code, and Codex. It routes each request to the cheapest model capable of doing the job, cutting inference spend automatically - and it is free while in beta.

Cloudflare Auto Router Cuts AI Inference Spend by Up to 30%

Cloudflare is turning its AI Gateway into an intelligent cost-control plane. Today the company released Auto Router in public beta, a feature that automatically routes every AI request to the most cost-efficient model capable of handling it, instead of leaving model selection to individual users reaching for a frontier model by default.

What happened

Announced as part of Cloudflare's Birthday Week, Auto Router lets developers set their model to cloudflare/auto inside AI Gateway. From there, the router selects a model for each request automatically. Cloudflare reports internal savings of up to 30 percent compared with sending every task to frontier models such as OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5, according to the announcement.

In Cloudflare's own general knowledge-work benchmark, the router completed 252 of 291 trials (86.6 percent) at 80 percent of the cost of GPT-6 Sol and 35 percent of the cost of Claude Opus 5.5. The core argument: not every task needs a frontier model, and paying frontier rates for non-frontier work is where most AI spend quietly leaks.

How it works

When a request hits cloudflare/auto, AI Gateway first builds a pool of models that can actually serve it, filtering out providers that do not support the request format, lack credentials, or are mid-outage. The remaining candidates are scored by a multi-head classification model running on Workers AI across Cloudflare's edge network, which reads the most recent turns of a conversation and predicts which model is capable enough for the job.

A model-routing hub distributing requests across a network

A key insight from Cloudflare: lower per-token prices do not always mean lower total cost. A cheap model that needs far more tokens to reach a correct answer can end up more expensive than a pricier one that finishes sooner. Auto Router optimizes for predicted trajectory cost, not dollars per million tokens.

Why it matters for builders

For teams shipping agents and internal AI tools, Auto Router removes a decision users currently make by hand inside harnesses such as OpenCode, Claude Code, and Codex. Budgets and spend limits only go so far, Cloudflare notes - the best savings are the ones users never notice. Auto Router is free while in beta, and Cloudflare frames it as the first step toward making the gateway itself act as a control plane on a user's behalf.

The broader signal is that model routing is consolidating into a standard infrastructure layer. OpenRouter and LiteLLM have shipped their own auto-routers, and now a major CDN sitting in the inference path is doing the same at the gateway level.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

30 September 2026

Updated

30 September 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.