If you plan to move Opus-pinned agents to Claude Sonnet 5.5 because it costs half as much per token, check the per-task numbers first. On Artificial Analysis's independent index, Sonnet 5.5 at max effort ties Opus 5.5 at xhigh effort, and costs $7.60 per task against Opus's $3.46. The cheaper per-token model costs more than twice as much per task. Sonnet 5.5 is the right call at low and high effort, where it has no Opus equivalent at its price. Once you push it far enough to match Opus, it stops saving money.
Anthropic released Claude Sonnet 5.5 on September 28 at the same $2 input / $10 output per million tokens as Sonnet 5. Opus 5.5 lists at $4 / $20. The launch chart places the two within a few points of each other, and the obvious response is to downgrade. That response is only right if you fix the effort level first.
What Anthropic's Launch Chart Actually Compares
Anthropic's chart does not hold effort constant across the two models. The launch page reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 against 66.4% for Opus 5.5 and 10.3% for Sonnet 5. It is 1844 against 1846 on GDPval-AA v2.1 and 80.1% against 81.8% on OSWorld 2.1. The footnotes say Opus 5.5's Terminal-Bench score is its xhigh result, while Sonnet 5.5's FrontierCode score is its max result.
Effort is Anthropic's per-request control over how many tokens the model may spend thinking, calling tools and writing. The effort documentation lists five levels, low through max, and says effort "affects all tokens in the response". Comparing two models at different effort levels is really a comparison of two token budgets.
FrontierCode leaves a real gap. Sonnet 5.5 scores 46.2% at max, per Anthropic's own table, against 54.4% for Opus 5.5, and a footnote concedes Sonnet scores lower at max than at xhigh. The Decoder's write-up notes that Anthropic's cost and speed claims still need independent testing to confirm them. The GDPval-AA and AA-Briefcase runs are the exception: Artificial Analysis ran those on a pre-release deployment.
Where Sonnet 5.5 Stops Being the Cheap Model
Measured per task instead of per token, Sonnet 5.5 is cheaper than Opus 5.5 only at the bottom of the effort range. Artificial Analysis ran both models at all five effort levels and published the score and cost to run its Intelligence Index for each one. The figures below come from its Sonnet 5.5 release page and Opus 5.5 release page, read 2026-09-29.
| Effort | Sonnet 5.5 score | Sonnet 5.5 $/task | Opus 5.5 score | Opus 5.5 $/task |
|---|---|---|---|---|
| low | 36 | $0.41 | 42 | $0.55 |
| medium | 41 | $0.59 | 51 | $1.34 |
| high | 47 | $1.08 | 54 | $1.82 |
| xhigh | 52 | $2.74 | 56 | $3.46 |
| max | 56 | $7.60 | 58 | $5.98 |
Compare them at matched scores:
- Sonnet 5.5 max (56, $7.60) vs Opus 5.5 xhigh (56, $3.46). The two score the same, and Sonnet costs 2.2x as much.
- Sonnet 5.5 xhigh (52, $2.74) vs Opus 5.5 high (54, $1.82). Opus scores higher and costs about a third less.
- Sonnet 5.5 medium (41, $0.59) vs Opus 5.5 low (42, $0.55). Opus costs about the same and scores one point higher.
The cause is token volume. Artificial Analysis found that Sonnet 5.5 at max used about 193,000 output tokens per index task, "the highest token use we have measured, around 60% higher than Opus 5.5 (max)". Output alone does not explain the gap: 60% more output tokens at half the rate is still about 20% cheaper. The rest of the bill comes from the work around those tokens, and Artificial Analysis does not publish that breakdown. What it does publish is the result: at the same score, Sonnet at max costs more than twice as much as Opus at xhigh.
Sonnet 5.5 has two points on the curve that Opus cannot reach. Low effort (36 for $0.41) is the cheapest point on either curve. High effort (47 for $1.08) sits between Opus low and Opus medium, and nothing from Opus lands there. If your workload is a well-specified subagent, a classifier step or high-volume chat, those are real savings. If it is a hard, long agentic task, the index says Opus at lower effort does the same work for less money.
One caveat applies to all of this. It is one composite index, and your workload is not that index. Artificial Analysis's own Terminal-Bench 4.0 run has Sonnet 5.5 ahead at 64% against 60% for Opus 5.5, so on pure terminal-driven coding the ranking can flip. That is why you need a bake-off, not a reason to skip one.
The Default Settings Skew Your Bake-Off
If you A/B the two models without setting effort, you will compare Sonnet at high against Opus at medium. The effort docs say Opus 5.5 "defaults to medium", while Sonnet 5.5 keeps high as its API default. The Sonnet 5.5 guide adds that its effort levels "are recalibrated", so a level does not produce the same amount of thinking it did on Sonnet 5.
Read against the index table, a default-versus-default test puts Sonnet 5.5 at 47 for $1.08 against Opus 5.5 at 51 for $1.34. Sonnet looks about 20% cheaper and a bit worse. That result tells you about the defaults, not the models. We made the same point about GPT-6 Sol's launch benchmarks at top effort and Grok 4.7's unchanged price. Per-token price is an input. Cost per task is what you actually pay.
Caching narrows the gap further. Opus 5.5 cache hits are billed at 0.05x base input, or $0.20 per million tokens, the same $0.20 Sonnet 5.5 charges. On a long-running agent that rereads a large, stable prefix every turn, much of the input bill costs the same on either model. What remains is uncached input and output, and Sonnet's token volume at high effort erodes its half-price rate on both. We worked through this caching crossover when Anthropic cut cache reads on Fable 5.1.
The steel-man for Sonnet is speed. Artificial Analysis measured Sonnet 5.5 at 142 output tokens per second at max, against 95 for Opus 5.5 at max. Anthropic quotes Zendesk's director of AI saying tickets were processed 20% faster. For an interactive agent where a human waits on every turn, the latency may be worth a higher bill. Make that trade on purpose, knowing its cost.
The Five Breaking Changes Behind the Model ID Swap
Swapping claude-sonnet-5 for claude-sonnet-5-5 is not a config change. Five request shapes that work on Sonnet 5 return a 400 on Sonnet 5.5. Anthropic's what's-new page lists them:
thinking: {"type": "disabled"}is rejected. Sendthinking: {"type": "between_tools"}instead. It is the lowest thinking setting, and it is only accepted at low, medium and high effort. At xhigh or max it also returns a 400. Withbetween_tools, effort cannot change mid-conversation.- Forced tool use is gone.
tool_choiceofanyortoolreturns an error. Useautowith strict tool use, or move the schema to structured outputs. Opus 5.5 made the same break the week before. - Thinking blocks are bound to the conversation. For accounts created on or after August 31, 2026, replaying a Sonnet 5.5 thinking block after editing the system prompt, the tools or an earlier message returns a 400.
computer_20251124is rejected on the Claude API and Google Cloud. Claude Computer Use integrations must move tocomputer_toolset_20260801, while Amazon Bedrock still accepts the old tool.- Opus 4.8, Opus 4.7 and Sonnet 5 advisors are rejected by a Sonnet 5.5 executor in the advisor tool.
The third one has already broken a real framework. An open issue on the Turnstone agent framework reports that claude-sonnet-5-5 fell through to Sonnet 5's capability settings. The first request after auto-compaction rewrote history then failed with a 400. Any harness that compacts, summarises or rewrites the transcript, which describes most long-running agents, needs the drop_block setting or append-only history before the swap.
A quieter change fails nothing and still breaks the product. Text the model writes between tool calls now comes back inside thinking blocks, empty at the default display setting. An agent UI that streams those progress notes to users goes silent between tool calls, and no error is raised.
What to Do Before You Move a Workload
Treat Sonnet 5.5 as a new model on its own effort curve, not a cheaper Opus. Anthropic's own developer guidance says the same thing: "for the hardest long-horizon work, an Opus model is the better choice."
This Week:
- Grep your codebase for
"disabled",tool_choiceandcomputer_20251124. Each hit is a 400 waiting behind the model swap. Anthropic's migration guide has a per-model checklist. - Find out when your API account was created. If it was on or after August 31, 2026, the thinking-block binding check is on by default. Check whether your agent framework compacts history.
- Set
effortexplicitly on every request to both models. Relying on the defaults means your test compares Sonnet high against Opus medium.
This Month:
- Run a matched-score bake-off, not a matched-effort one. Take 50 to 100 real tasks from your traffic. Run Sonnet 5.5 at low, medium and high, and Opus 5.5 at low and medium. Record pass rate, total tokens and dollars per completed task. Where two settings pass the same share of tasks, the cheaper one wins. Our piece on harness design flipping model rankings explains why this has to run on your harness, not a vendor's.
- Route by effort band, not by model name. Put high-volume, well-specified steps on Sonnet 5.5 low. Put the hard, long-horizon work on Opus 5.5 at medium or high, and stop running anything at max until an eval shows it pays. A gateway such as LiteLLM can hold both routes. Our model router buyer's guide covers what to buy and what to build.
Before Renewal:
- Price your commit in dollars per completed task, not in tokens. A committed-spend deal sized on Sonnet's per-token rate will run short if a migration moves hard work onto Sonnet at max effort.
The Bottom Line
Sonnet 5.5 is a real improvement on Sonnet 5, and it does not replace Opus 5.5 for hard agentic work. In July, Sonnet 5 launched with the same pitch, Opus performance at a fraction of the price. The pitch holds at the effort level the vendor chose for the chart. It fails at the effort level your hardest workload needs.
This is the pattern of the reasoning-model era. Per-token price went from a bill to a coefficient, and effort is the variable it multiplies. Two models can no longer be compared on a price sheet. They can only be compared on your tasks, at matched quality, in dollars.
Half the token price buys you nothing if the model needs max effort to finish the job.
Continue Reading
- Claude Opus 5.5 Rejects the Forced Tool Calls Opus 5 Accepted
- GPT-6 Sol Halves the Token Price but Benchmarks It at Top Effort
- Grok 4.7 Kept Its $2/$6 Price and Can Still Double Cost per Task
- Model Router Buyer's Guide: Buy Failover, Not Judgment
- Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.
- Your Bake-Off Ranked Harnesses, Not Models. Re-Run It.
