Cloudflare's new Auto Router is cheaper per solved task than Claude Opus 5.5 and fails nearly four times as often. On Cloudflare's own 97-task benchmark, cloudflare/auto succeeded 86.6% of the time at $0.0084 per success; Opus 5.5 succeeded 96.6% of the time at $0.0210, per Cloudflare's launch post. Run the arithmetic and the router only saves money if a failed task costs your team less than about 13 cents to notice and redo. For most agent workloads it costs far more. Automatic routing is a quality decision presented as a cost one, and you should make it per workload rather than per gateway.
That doesn't make the product bad. It is the most transparent router launch so far: Cloudflare published the classifier design, the cache logic and an eval with the losing number in it. That openness is why you can do the math before you flip the switch.
What Cloudflare Actually Shipped
Auto Router is a model name, cloudflare/auto, that you send to AI Gateway in place of a specific provider model, and the gateway picks the model. It launched in public beta on September 30, 2026, and it is free while in beta.
According to the developer documentation, the default pool covers three providers: anthropic/claude-fable-5, anthropic/claude-opus-5, anthropic/claude-sonnet-5, openai/gpt-5.6-luna, openai/gpt-5.6-sol, openai/gpt-5.6-terra and xai/grok-4.5. Two request headers narrow it. cf-aig-allowed-models takes a comma-separated list with wildcards such as anthropic/*, and it replaces the defaults. cf-aig-allowed-providers restricts by vendor. Every response comes back with three routing headers: cf-aig-routed-model, cf-aig-routing-reason and cf-aig-routing-decision-id.
The routing-reason values are worth reading. Besides cost_optimal_within_pool and pinned_by_turn, the docs list fallback_router_error, fallback_router_timeout, fallback_candidate_unavailable and fallback_unsupported_input. Put plainly, sometimes the router itself fails and the gateway falls back. You should log how often that happens.
The documents disagree on one point. The docs say Auto Router "supports the Chat Completions and Responses API formats." The blog lists "full support for the Responses API and WebSockets" as future work. If your agents run on the Responses API, test that path before you trust it.
How Does Auto Router Decide Which Model Gets a Prompt?
It scores your request with a small classifier, not with an LLM call. Cloudflare describes a multi-head classification model running on Workers AI on GPUs across its edge network. The model does two things.
First, it assigns probabilities across 14 task categories. Cloudflare names a few (coding, planning, research, data analysis) but does not publish the full list. Second, it rates the request from 1 to 5 on four dimensions: complexity, ambiguity, stakes, and how much the request depends on earlier context. The router drops models that can't handle the request format or execution mode, takes the gateway's credentials, billing configuration, access policies and spend limits into account, and picks the cost-optimal model from what remains. If that provider can't serve the request, the gateway "can move to another eligible model."
A model router is a layer that decides, request by request, which model answers. That decision used to sit in your code as a hard-coded model string. The design choice that matters here is "stakes" as an input dimension. Cloudflare is admitting that cost-optimal routing depends on how bad a wrong answer would be, and its classifier has to guess that from the prompt text. You know it from the business process. That gap is the subject of the rest of this piece.
Why the Router Refuses to Switch Mid-Conversation
Auto Router treats switching models as a cost, because it is one. Cloudflare's rule within a turn is that "switching rarely pays off, so it's better to keep using the same model." Across turns, a switching penalty grows with the number of tokens already in context: "the deeper the conversation, the more a switch has to earn back."
The reason is prompt caching. A prompt cache stores the processed prefix of a conversation so the next call re-reads it cheaply instead of re-processing it at full price. Anthropic charges cache reads at 0.1x the base input price, and 0.05x on Opus 5.5, while a 5-minute cache write costs 1.25x. OpenAI bills cached input at 0.1x uncached input (0.05x on GPT-6.1 Sol) and turns caching on by default. A cache belongs to one model's prefix, so moving a 100,000-token agent session to another model means paying to write the whole context again, at 12.5 to 25 times the Anthropic read price.
Cloudflare names a second, less obvious cost: "most models can't read another model's reasoning tokens," so a switch can force the new model to redo that thinking at output prices. To keep a conversation on one model, send cf-aig-session-id; the docs say the router then switches at the start of a turn "only when the expected benefit outweighs the cost of losing the cache."
We made this argument in August, before this product existed: on agentic workloads, a naive router gives up a 10x caching discount to chase a smaller per-token one. Cloudflare built that penalty into the router. That's the right engineering. It also means that in long agent sessions the router mostly chooses at turn one and then stays put, so the turn-one classification carries most of the decision.
What Cloudflare's Own Benchmark Shows
The router lands between the two frontier models on quality and beats both on cost per success. The benchmark uses simulated workspace tools across email, calendars, Slack, files, travel and finance, and every task needs "a verifiable answer or complete an action." It has 97 tasks, sampled three times per model, for 291 trials each.
| Model | Success rate | Cost per success | Failures per 1,000 tasks |
|---|---|---|---|
| Claude Opus 5.5 | 96.6% | $0.0210 | 34 |
cloudflare/auto |
86.6% | $0.0084 | 134 |
| GPT-6 Sol | 84.2% | $0.0108 | 158 |
Success and cost from Cloudflare's launch post, 2026-09-30. Failures per 1,000 derived from the success rate.
Cloudflare summarises the result as "similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus." Against Sol, that's a fair reading. The 2.4-point lead amounts to about seven trials out of 291, which is within noise for a set this size. Small agent benchmarks rarely separate models that close. Against Opus, "similar" does not describe a 10-point gap: 134 failures per thousand against 34, or 3.9 times as many.
Note what the table does not tell you. The default pool in the docs lists claude-opus-5 and gpt-5.6-sol, and includes neither Opus 5.5 nor GPT-6 Sol, the two models the router is benchmarked against. The post does not say which candidate pool the router drew from in the eval. The separate claim of "cost savings of up to 30%" comes from Cloudflare running the router internally through its OpenCode harness, with no task count or success rate attached. Both numbers are vendor claims and should be weighed as such.
When Does Opus Beat the Router on Cost?
Opus 5.5 is cheaper whenever a failed task costs more than about 13 cents to catch and redo. Here's the derivation. Cost per success is total spend divided by successes, so total spend is cost per success times successes:
- Router, per 1,000 tasks: 866 successes × $0.0084 ≈ $7.27
- Opus 5.5, per 1,000 tasks: 966 successes × $0.0210 ≈ $20.29
The router saves about $13.02 per thousand tasks and fails about 100 more times. That's $0.13 saved per extra failure. Any failure that takes a person ten seconds to spot, or that triggers a retry, an escalation or a wrong calendar invite to a customer, costs more than that.
The strongest case for the router: plenty of traffic really does have near-zero failure cost. Examples are drafts a human edits anyway, internal summaries, classification with a downstream check, and chat where users simply ask again. On those workloads, cutting the bill to 35% of Opus is real money, and a router is a better tool than a hard-coded cheap model because it moves hard prompts up a tier on its own.
The mistake is applying one gateway-wide setting across both kinds of traffic. The router's "stakes" dimension is its guess at what you already know from the business process. Our router buyer's guide reached the same conclusion across Bedrock and Azure: buy the failover, and keep the judgement about which workload gets which tier for yourself.
What Data-Retention Terms Does Routed Traffic Get?
Today, whatever terms the selected model's path carries. Auto Router does not filter on them yet. Cloudflare's future-work list includes "include zero-data-retention requirements when filtering models", so that filter does not exist in the beta.
AI Gateway does offer zero data retention elsewhere. Under Unified Billing, "ZDR routes Unified Billing traffic through provider endpoints that do not retain prompts or responses," but only for requests using Cloudflare-managed credentials, not bring-your-own-key. Cloudflare tells you to check the model catalog for which models support it. Unified Billing also adds a 5% fee on purchased credits.
For a regulated team, that settles it. If your data-processing terms name specific providers or require ZDR, cf-aig-allowed-models is a compliance control, not a tuning knob. Pin it to the models whose retention terms you have actually reviewed. An unpinned default pool is a list Cloudflare can expand. "Expand models offered through cloudflare/auto" is the first item on the same roadmap. When Stripe bought OpenRouter we made a similar point: a gateway setting is not a contract.
Haven't We Tried Automatic Routing Before?
Yes, and the earlier claims sounded similar. The 2024 RouteLLM paper from the LMSYS team reported cutting costs "by over 2 times in certain cases" without compromising quality, by routing between one strong and one weak model on public benchmarks. OpenRouter's openrouter/auto classifies prompts into roughly 30 task categories and ranks models by what OpenRouter's users spend on each, with no added fee and an allowed_models filter that uses the same wildcard syntax.
Cloudflare's design is more cautious than either: a cache-aware switching penalty and a published loss against the frontier model. The pattern repeats, though. Routers report cost savings on benchmarks where everything is graded the same, and buyers find out that their own failures are not. Cloudflare also sells a deterministic option. Dynamic routing lets you write the rules yourself: conditionals on request metadata, percentage splits, and budget-limit nodes that fall back when a key hits its cap. For most enterprises, that plus Auto Router on a limited allowed list is the right combination.
What to Do Before You Set Model to cloudflare/auto
This Week:
- Sort your gateway traffic by failure cost, not by volume. Write down, per workload, what one wrong answer costs: seconds of review, a retry, or a customer-facing error. Anything above the 13-cent line stays pinned to a named model.
- Pin
cf-aig-allowed-modelsbefore the first routed request. List only the models whose retention and processing terms your legal team has reviewed. Do not run on the default pool. - Log all three response headers. Store
cf-aig-routed-model,cf-aig-routing-reasonandcf-aig-routing-decision-idwith each trace, so a bad answer can be tied to the model that produced it, and count how often you seefallback_router_*.
This Month:
- Shadow-route one low-stakes workload. Send it to
cloudflare/autowithcf-aig-session-idset, keep your pinned model as the control, and compare success rate on your own graded set, not Cloudflare's 97 tasks. - Measure the cache effect on your longest sessions. Compare cache-read share before and after routing. If it falls, the switching penalty isn't protecting you on that workload.
- Write the routing rules you already know into dynamic routing. Use deterministic routes for the high-stakes paths and Auto Router only inside the low-stakes branch.
Before the Beta Ends:
- Get the post-beta price and the ZDR filter date in writing. "Free while in beta" has an end date. Budget for the router and for the 5% Unified Billing fee if you use managed credentials.
The Bottom Line
Cloudflare has shipped the most honest router launch yet: a stated classifier, a real answer to the caching problem, and a benchmark that shows its own product 10 points behind Opus. The previous wave of routers promised savings without losing quality. This one shows the quality loss, and that lets you decide where it is acceptable.
That decision belongs to you. A classifier can estimate "stakes" from a prompt. Only you know what a failed task costs. Route the cheap failures and pin the expensive ones.
Continue Reading
- Your AI Router Is Trading a 10x Discount for a 2.5x One
- Model Router Buyer's Guide: Buy Failover, Not Judgment
- Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- Best LLM Gateways for Cost Control: Self-Host First
- Stripe Bought OpenRouter. A Toggle Is Not a Contract.
- Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score
