Cloudflare has released Auto Router in public beta through AI Gateway. Setting a model to cloudflare/auto routes each request to a model judged capable enough for the task, so users do not choose a model themselves. Cloudflare reports early internal results through its OpenCode harness showing cost savings of up to 30% versus using only frontier models such as OpenAI Sol and Anthropic Claude Opus.
From visibility to automatic decisions
Cloudflare's earlier work on AI spend covered setting budgets and limits, and linking employees to their usage. Those controls give organizations guardrails, but they still depend on individuals making cost-conscious choices request by request. In harnesses including OpenCode, Claude Code, and Codex, users select models manually, and tasks of unequal difficulty often get served by models that are overkill — Opus-level intelligence is not needed to summarize an email, though blocking that model entirely would hurt a security engineering team.
AI Gateway sits in the path of every request from every user, agent, and tool, which Cloudflare frames as the reason it can go beyond observing and enforcing. The stated goal is for the gateway to make routing decisions on the user's behalf, reducing spend automatically while preserving access to the most capable models when work requires them.
Measured results
Cloudflare uses Auto Router internally in its OpenCode deployment and in Cloudflare OS, its custom agent harness. Internal usage produced results comparable with frontier models on coding tasks. The router performs best across a broad range of knowledge-work tasks, such as those spread across technical and non-technical teams in a large organization.
Cloudflare evaluated cloudflare/auto against OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 on an internal general knowledge work benchmark. The benchmark uses simulated workspace tools and spans day-to-day workflows in email, calendars, Slack, files, travel and finance, with each task requiring tool use to produce a verifiable answer or complete an action.
Model | Successful Trials | Success Rate | Total Cost | Cost per success |
cloudflare/auto | 252/291 | 86.6% (+6.2/−6.9 pp) | $2.10 | $0.0084 |
Anthropic Claude Opus 5.5 | 281/291 | 96.6% (+2.7/−3.8 pp) | $5.91 | $0.0210 |
OpenAI GPT-6 Sol | 245/291 | 84.2% (+6.5/−6.9 pp) | $2.64 | $0.0108 |
Performance was similar to other state-of-the-art daily-driver models, at 80% the cost of Sol and 35% the cost of Opus. Cloudflare attributes this to the "jagged frontier" across models: the ability to solve a given problem generally exists somewhere in the portfolio, and the router's job is to pick the right one while balancing quality and price. Savings come from not paying frontier rates for non-frontier work, and they scale with how much of that work exists.
Cloudflare also notes that lower token prices do not always produce lower-cost outcomes — a model that looks cheaper per token may consume disproportionately more tokens to solve a problem. A router should therefore minimize predicted trajectory cost rather than load-balance on dollars per million tokens.
How routing works

When a request arrives at cloudflare/auto, AI Gateway first assembles the pool of models that can serve it. Models that do not support the request format or execution mode are filtered out, along with consideration of credentials, billing configuration, access control policies, and spend limits attached to the gateway. Unhealthy upstream providers or models are excluded during downtime and returned to the pool after an outage.
For the remaining candidates, the router examines a compact view of the conversation, prioritizing the newest turns. That conversation goes to a multi-head classification model running on Workers AI, deployed on GPUs across Cloudflare's edge network. The classifier outputs two sets of signals: probabilities across 14 task categories such as coding, planning, research and data analysis, and ratings across four dimensions on a one-to-five scale — complexity, ambiguity, stakes, and dependence on earlier context.
A separate scoring matrix combines those signals with model benchmark results to estimate fit. It was calibrated by defining the preferred model for a set of example task and difficulty profiles, then adjusting the weights to reproduce those choices. The router then combines expected quality with each model's input and output token prices: on straightforward requests price carries more weight, so a smaller model can win when it is capable enough, while rising difficulty reduces the cost penalty and gives stronger models more room. In simplified terms, it selects the model with the highest utility as defined by:
utility = expected quality - adaptive cost penalty
Long agentic sessions such as debugging or coding are driven less by list price than by cache reads, which grow with session length. Switching models discards the cache and forces the new model to write the whole context again — a cost that can be worth paying if a model's cache-read and cache-write prices are cheap enough to pay it back quickly.
Rather than avoiding switching altogether, the Auto Router prices cache reads and writes into its decision. Within a turn, meaning one user input loop, the cache is hot and switching rarely pays off, so staying on the same model is preferable. Across turns, a switching penalty grows with the number of tokens already in context: a model still holding a live cache for the session is priced at its cheaper cache-read rate, while every other candidate is priced at the full cost of rewriting the context. The deeper the conversation, the more a switch must earn back through higher quality, fewer tokens overall, or cheaper cache rereads. Model switching carries a second cost, since most models cannot read another model's reasoning tokens, so dropping them may force the new model to redo that work at output prices. Cloudflare wants the router to prefer staying within the same model family when it does switch.
The router returns a ranked list; AI Gateway attempts the winner first and can fall back to another eligible model if that provider cannot serve the request.
Cloudflare points to several benefits of the two-stage design, from classification to scoring matrix. Routing decisions are legible, because the predicted category and complexity of a task can be inspected alongside the resulting model choice. Adding a newly released model requires no retraining — only its benchmark-derived weights in the scoring matrix. The same classifier can also back different routing profiles: alongside cloudflare/auto, Cloudflare plans other routers including cloudflare/auto-best, which uses the same classification and model pool but selects the highest expected quality without applying the cost tradeoff.
Roadmap and availability
Planned near-term work includes expanding the models offered through cloudflare/auto, factoring zero-data-retention requirements into model filtering, accounting for provider capacity during selection, choosing the appropriate reasoning or thinking level per request, adding full support for the Responses API and WebSockets, and exploring structured decision models as a first-pass classifier.
The Auto Router is free while in beta, with further detail in Cloudflare's developer documentation.



