Cloudflare's Auto Router Fails 4x as Often as Opus on Its Own Test

Cloudflare's Auto Router solved 86.6% of its own benchmark tasks against Opus 5.5's 96.6%, at 40% of the cost per success. The saving only holds where a failed task costs under about 13 cents.

By Rajesh Beri·October 1, 2026·11 min read
Share:
A railway signal box interior with a single operator's lever frame, one heavy lever pulled toward a track labelled only by coloured lamps, a rack of server blades glowing behind the levers.

Illustration generated using AI

Cloudflare's new Auto Router is cheaper per solved task than Claude Opus 5.5 and fails nearly four times as often. On Cloudflare's own 97-task benchmark, cloudflare/auto succeeded 86.6% of the time at $0.0084 per success; Opus 5.5 succeeded 96.6% of the time at $0.0210, per Cloudflare's launch post. Run the arithmetic and the router only saves money if a failed task costs your team less than about 13 cents to notice and redo. For most agent workloads it costs far more. Automatic routing is a quality decision presented as a cost one, and you should make it per workload rather than per gateway.

That doesn't make the product bad. It is the most transparent router launch so far: Cloudflare published the classifier design, the cache logic and an eval with the losing number in it. That openness is why you can do the math before you flip the switch.

What Cloudflare Actually Shipped

Auto Router is a model name, cloudflare/auto, that you send to AI Gateway in place of a specific provider model, and the gateway picks the model. It launched in public beta on September 30, 2026, and it is free while in beta.

According to the developer documentation, the default pool covers three providers: anthropic/claude-fable-5, anthropic/claude-opus-5, anthropic/claude-sonnet-5, openai/gpt-5.6-luna, openai/gpt-5.6-sol, openai/gpt-5.6-terra and xai/grok-4.5. Two request headers narrow it. cf-aig-allowed-models takes a comma-separated list with wildcards such as anthropic/*, and it replaces the defaults. cf-aig-allowed-providers restricts by vendor. Every response comes back with three routing headers: cf-aig-routed-model, cf-aig-routing-reason and cf-aig-routing-decision-id.

The routing-reason values are worth reading. Besides cost_optimal_within_pool and pinned_by_turn, the docs list fallback_router_error, fallback_router_timeout, fallback_candidate_unavailable and fallback_unsupported_input. Put plainly, sometimes the router itself fails and the gateway falls back. You should log how often that happens.

The documents disagree on one point. The docs say Auto Router "supports the Chat Completions and Responses API formats." The blog lists "full support for the Responses API and WebSockets" as future work. If your agents run on the Responses API, test that path before you trust it.

How Does Auto Router Decide Which Model Gets a Prompt?

It scores your request with a small classifier, not with an LLM call. Cloudflare describes a multi-head classification model running on Workers AI on GPUs across its edge network. The model does two things.

First, it assigns probabilities across 14 task categories. Cloudflare names a few (coding, planning, research, data analysis) but does not publish the full list. Second, it rates the request from 1 to 5 on four dimensions: complexity, ambiguity, stakes, and how much the request depends on earlier context. The router drops models that can't handle the request format or execution mode, takes the gateway's credentials, billing configuration, access policies and spend limits into account, and picks the cost-optimal model from what remains. If that provider can't serve the request, the gateway "can move to another eligible model."

A model router is a layer that decides, request by request, which model answers. That decision used to sit in your code as a hard-coded model string. The design choice that matters here is "stakes" as an input dimension. Cloudflare is admitting that cost-optimal routing depends on how bad a wrong answer would be, and its classifier has to guess that from the prompt text. You know it from the business process. That gap is the subject of the rest of this piece.


Why the Router Refuses to Switch Mid-Conversation

Auto Router treats switching models as a cost, because it is one. Cloudflare's rule within a turn is that "switching rarely pays off, so it's better to keep using the same model." Across turns, a switching penalty grows with the number of tokens already in context: "the deeper the conversation, the more a switch has to earn back."

The reason is prompt caching. A prompt cache stores the processed prefix of a conversation so the next call re-reads it cheaply instead of re-processing it at full price. Anthropic charges cache reads at 0.1x the base input price, and 0.05x on Opus 5.5, while a 5-minute cache write costs 1.25x. OpenAI bills cached input at 0.1x uncached input (0.05x on GPT-6.1 Sol) and turns caching on by default. A cache belongs to one model's prefix, so moving a 100,000-token agent session to another model means paying to write the whole context again, at 12.5 to 25 times the Anthropic read price.

Cloudflare names a second, less obvious cost: "most models can't read another model's reasoning tokens," so a switch can force the new model to redo that thinking at output prices. To keep a conversation on one model, send cf-aig-session-id; the docs say the router then switches at the start of a turn "only when the expected benefit outweighs the cost of losing the cache."

We made this argument in August, before this product existed: on agentic workloads, a naive router gives up a 10x caching discount to chase a smaller per-token one. Cloudflare built that penalty into the router. That's the right engineering. It also means that in long agent sessions the router mostly chooses at turn one and then stays put, so the turn-one classification carries most of the decision.

What Cloudflare's Own Benchmark Shows

The router lands between the two frontier models on quality and beats both on cost per success. The benchmark uses simulated workspace tools across email, calendars, Slack, files, travel and finance, and every task needs "a verifiable answer or complete an action." It has 97 tasks, sampled three times per model, for 291 trials each.

Model Success rate Cost per success Failures per 1,000 tasks
Claude Opus 5.5 96.6% $0.0210 34
cloudflare/auto 86.6% $0.0084 134
GPT-6 Sol 84.2% $0.0108 158

Success and cost from Cloudflare's launch post, 2026-09-30. Failures per 1,000 derived from the success rate.

Cloudflare summarises the result as "similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus." Against Sol, that's a fair reading. The 2.4-point lead amounts to about seven trials out of 291, which is within noise for a set this size. Small agent benchmarks rarely separate models that close. Against Opus, "similar" does not describe a 10-point gap: 134 failures per thousand against 34, or 3.9 times as many.

Note what the table does not tell you. The default pool in the docs lists claude-opus-5 and gpt-5.6-sol, and includes neither Opus 5.5 nor GPT-6 Sol, the two models the router is benchmarked against. The post does not say which candidate pool the router drew from in the eval. The separate claim of "cost savings of up to 30%" comes from Cloudflare running the router internally through its OpenCode harness, with no task count or success rate attached. Both numbers are vendor claims and should be weighed as such.

When Does Opus Beat the Router on Cost?

Opus 5.5 is cheaper whenever a failed task costs more than about 13 cents to catch and redo. Here's the derivation. Cost per success is total spend divided by successes, so total spend is cost per success times successes:

  • Router, per 1,000 tasks: 866 successes × $0.0084 ≈ $7.27
  • Opus 5.5, per 1,000 tasks: 966 successes × $0.0210 ≈ $20.29

The router saves about $13.02 per thousand tasks and fails about 100 more times. That's $0.13 saved per extra failure. Any failure that takes a person ten seconds to spot, or that triggers a retry, an escalation or a wrong calendar invite to a customer, costs more than that.

The strongest case for the router: plenty of traffic really does have near-zero failure cost. Examples are drafts a human edits anyway, internal summaries, classification with a downstream check, and chat where users simply ask again. On those workloads, cutting the bill to 35% of Opus is real money, and a router is a better tool than a hard-coded cheap model because it moves hard prompts up a tier on its own.

The mistake is applying one gateway-wide setting across both kinds of traffic. The router's "stakes" dimension is its guess at what you already know from the business process. Our router buyer's guide reached the same conclusion across Bedrock and Azure: buy the failover, and keep the judgement about which workload gets which tier for yourself.


What Data-Retention Terms Does Routed Traffic Get?

Today, whatever terms the selected model's path carries. Auto Router does not filter on them yet. Cloudflare's future-work list includes "include zero-data-retention requirements when filtering models", so that filter does not exist in the beta.

AI Gateway does offer zero data retention elsewhere. Under Unified Billing, "ZDR routes Unified Billing traffic through provider endpoints that do not retain prompts or responses," but only for requests using Cloudflare-managed credentials, not bring-your-own-key. Cloudflare tells you to check the model catalog for which models support it. Unified Billing also adds a 5% fee on purchased credits.

For a regulated team, that settles it. If your data-processing terms name specific providers or require ZDR, cf-aig-allowed-models is a compliance control, not a tuning knob. Pin it to the models whose retention terms you have actually reviewed. An unpinned default pool is a list Cloudflare can expand. "Expand models offered through cloudflare/auto" is the first item on the same roadmap. When Stripe bought OpenRouter we made a similar point: a gateway setting is not a contract.

Haven't We Tried Automatic Routing Before?

Yes, and the earlier claims sounded similar. The 2024 RouteLLM paper from the LMSYS team reported cutting costs "by over 2 times in certain cases" without compromising quality, by routing between one strong and one weak model on public benchmarks. OpenRouter's openrouter/auto classifies prompts into roughly 30 task categories and ranks models by what OpenRouter's users spend on each, with no added fee and an allowed_models filter that uses the same wildcard syntax.

Cloudflare's design is more cautious than either: a cache-aware switching penalty and a published loss against the frontier model. The pattern repeats, though. Routers report cost savings on benchmarks where everything is graded the same, and buyers find out that their own failures are not. Cloudflare also sells a deterministic option. Dynamic routing lets you write the rules yourself: conditionals on request metadata, percentage splits, and budget-limit nodes that fall back when a key hits its cap. For most enterprises, that plus Auto Router on a limited allowed list is the right combination.

What to Do Before You Set Model to cloudflare/auto

This Week:

  1. Sort your gateway traffic by failure cost, not by volume. Write down, per workload, what one wrong answer costs: seconds of review, a retry, or a customer-facing error. Anything above the 13-cent line stays pinned to a named model.
  2. Pin cf-aig-allowed-models before the first routed request. List only the models whose retention and processing terms your legal team has reviewed. Do not run on the default pool.
  3. Log all three response headers. Store cf-aig-routed-model, cf-aig-routing-reason and cf-aig-routing-decision-id with each trace, so a bad answer can be tied to the model that produced it, and count how often you see fallback_router_*.

This Month:

  1. Shadow-route one low-stakes workload. Send it to cloudflare/auto with cf-aig-session-id set, keep your pinned model as the control, and compare success rate on your own graded set, not Cloudflare's 97 tasks.
  2. Measure the cache effect on your longest sessions. Compare cache-read share before and after routing. If it falls, the switching penalty isn't protecting you on that workload.
  3. Write the routing rules you already know into dynamic routing. Use deterministic routes for the high-stakes paths and Auto Router only inside the low-stakes branch.

Before the Beta Ends:

  1. Get the post-beta price and the ZDR filter date in writing. "Free while in beta" has an end date. Budget for the router and for the 5% Unified Billing fee if you use managed credentials.

The Bottom Line

Cloudflare has shipped the most honest router launch yet: a stated classifier, a real answer to the caching problem, and a benchmark that shows its own product 10 points behind Opus. The previous wave of routers promised savings without losing quality. This one shows the quality loss, and that lets you decide where it is acceptable.

That decision belongs to you. A classifier can estimate "stakes" from a prompt. Only you know what a failed task costs. Route the cheap failures and pin the expensive ones.

Continue Reading

Share:

Frequently Asked Questions

What is Cloudflare Auto Router?

Auto Router is a model name, cloudflare/auto, that you send to Cloudflare AI Gateway instead of a specific model. A classifier running on Workers AI scores each request across 14 task categories and four 1-5 dimensions (complexity, ambiguity, stakes, context dependence) and picks the cost-optimal model from the allowed pool. It launched in public beta on September 30, 2026 and is free during the beta.

How does Cloudflare Auto Router compare to Claude Opus 5.5?

On Cloudflare's own 97-task workspace-tools benchmark (three samples per model), cloudflare/auto succeeded 86.6% of the time at $0.0084 per success, against Opus 5.5's 96.6% at $0.0210. GPT-6 Sol scored 84.2% at $0.0108. The router is cheaper per success but fails about 3.9 times as often as Opus.

Does Cloudflare Auto Router switch models mid-conversation?

Rarely. Within a turn it keeps the same model, and across turns it applies a switching penalty that grows with the tokens already in context, because switching loses the prompt cache and the previous model's reasoning tokens. Sending a cf-aig-session-id header keeps a conversation on one model unless a switch is expected to pay for itself.

Can I restrict which models Cloudflare Auto Router uses?

Yes. The cf-aig-allowed-models header takes a comma-separated list with wildcards such as anthropic/* and replaces the default pool, and cf-aig-allowed-providers restricts by vendor. Zero-data-retention filtering is not yet built into the router, so teams with retention requirements should pin the allowed list to models they have reviewed.

How do I see which model Auto Router picked?

Each response carries cf-aig-routed-model, cf-aig-routing-reason and cf-aig-routing-decision-id headers. Routing reasons include cost_optimal_within_pool, pinned_by_turn and several fallback values such as fallback_router_timeout, which are worth logging and counting.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →