Claude Haiku 5.5 Charges 5x Once a Prompt Passes 100K Tokens

Claude Haiku 5.5 matches GPT-6 Luna at $0.10/$0.50 per million tokens and beats it on Anthropic's evals, but prompts over 100K tokens cost 5x and its tokenizer counts about 30% more tokens than Haiku 4.5. Short-prompt classification and routing get the big saving; compaction and long-context work get much less.

By Rajesh Beri·October 7, 2026·11 min read
Share:
A server-room billing printer feeding out a long paper invoice on which a single line of figures abruptly jumps in size halfway down the page, with a small and a large stack of punched paper tapes beside it on a steel de

Illustration generated using AI

Claude Haiku 5.5 is about 87% cheaper than Haiku 4.5 for the same text if your prompts stay under 100,000 tokens, and about 35% cheaper if they don't. Anthropic released it on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output tokens, the same list price as OpenAI's GPT-6 Luna. A prompt over 100,000 tokens pays $0.50 and $2.50 instead, five times the short rate, and the model's tokenizer counts the same text as roughly 30% more tokens than Haiku 4.5 did. Your saving depends on how long your prompts are, and you can check that from logs you already have.

The benchmark story is real but narrow. On Anthropic's own table, Haiku 5.5 beats Luna on every eval where both have a score and loses to Claude Sonnet 5.5 on every one. All of those numbers are the vendor's.

What Does Claude Haiku 5.5 Cost?

Claude Haiku 5.5 has two price lists, split at a 100,000-token prompt. Anthropic's pricing page lists the short tier at $0.10 input, $0.50 output, $0.125 for a 5-minute cache write and $0.01 for a cache read, per million tokens. Over 100,000 tokens, every one of those rates goes up 5x: $0.50 input, $2.50 output, $0.625 cache write and $0.05 cache read. The Batch API halves both tiers, to $0.05/$0.25 short and $0.25/$1.25 long.

For comparison, Haiku 4.5 is $1 input and $5 output at any length, and Sonnet 5.5 is $2 and $10. Sonnet 5.5 has no length tier at all: the same pricing page says Claude 4.6 and later models, except Haiku 5.5, bill the full 1M-token context window at standard rates.

Per 1M tokens Haiku 5.5 (≤100K prompt) Haiku 5.5 (>100K prompt) Haiku 4.5 Sonnet 5.5 GPT-6 Luna (≤272K / >272K)
Input $0.10 $0.50 $1.00 $2.00 $0.10 / $0.20
Output $0.50 $2.50 $5.00 $10.00 $0.50 / $0.75
Cache read $0.01 $0.05 $0.10 $0.10 $0.01 / $0.02

Anthropic and OpenAI prices as listed on each vendor's pricing page on October 7, 2026.

Anthropic's launch post frames the cut as 90% off Haiku 4.5 for prompts up to 100,000 tokens and 50% off above it, and says about 90% of Haiku 4.5 requests fell in the short tier. It also says the model costs "around 75% less to run" on average. Anthropic doesn't break down the gap between 90% and 75%. The tokenizer and the long-prompt tail are the obvious candidates, and the published numbers are consistent with them, but Anthropic hasn't said so.

Why the Tokenizer Eats Part of the Discount

The same text costs more tokens on Haiku 5.5 than on Haiku 4.5. The Haiku 5.5 migration guide says the model uses the tokenizer introduced with Claude Opus 4.7, and "the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5," with the exact figure depending on content. It tells you to recount prompts and recompute cost estimates rather than reuse Haiku 4.5 counts.

Tokenizer inflation is the extra token count a newer tokenizer produces for identical text. Apply Anthropic's 30% to the list prices and the per-text discount moves:

  • Short prompts: $0.10 × 1.3 is about $0.13 per million of your old Haiku 4.5 tokens, against $1.00. Roughly 87% cheaper, not 90%.
  • Long prompts: $0.50 × 1.3 is about $0.65, against $1.00. Roughly 35% cheaper, not 50%.

The tokenizer also moves the threshold. A prompt that measured about 77,000 tokens on Haiku 4.5 comes out near 100,000 on Haiku 5.5, so some requests your dashboards file as "comfortably under 100K" will land in the expensive tier after the switch. Anthropic's launch post doesn't say whether its 90% short-tier share was counted with the old tokenizer or the new one.

The Decoder's launch coverage made the same point: real savings will be smaller than the per-token prices suggest. We worked through how tokenizer changes distort per-token comparisons in our inference cost guide.


Where the 100K Tier Hits Hardest

Haiku 5.5 is cheapest on exactly the work it is marketed for, except the work that reads long transcripts. Anthropic's launch post recommends it for "compaction, summarization, or subagent work." Classification, routing and extraction prompts are usually a few thousand tokens and stay in the cheap tier. Compaction is a summary of a long conversation, so by definition its prompt is long, and a compaction call on a 150,000-token agent transcript bills at the $0.50/$2.50 rate.

Sub-agents are the other trap. A sub-agent conversation grows every turn, and once its prompt passes 100,000 tokens every later turn pays the long rate. Cache reads don't escape: the Haiku 5.5 model page lists a separate cache-read price for each tier ($0.01 short, $0.05 long). The docs price by "prompt length" and don't spell out how cached prefix tokens count toward the 100K line, so send a test request just over the boundary and read the usage block before you model it.

A worked example, using OpenAI's and Anthropic's list prices and ignoring the tokenizer difference between vendors (which you should measure): a 150,000-token retrieval prompt with a 1,000-token answer.

  • Haiku 5.5: 150,000 × $0.50 + 1,000 × $2.50 per million, about $0.078.
  • GPT-6 Luna: under its 272K long-context line, 150,000 × $0.10 + 1,000 × $0.50 per million, about $0.016.

At that prompt size Luna is roughly 5x cheaper per token. Above 272K, Luna moves to $0.20 input and $0.75 output, and per a rate breakdown by eesel that rate applies to the whole request; Haiku 5.5's long tier is still 2.5x Luna on input and over 3x on output. The two vendors match on price only for short prompts. If your long-context work sits on Haiku today, our long context vs RAG cost analysis is the place to start.

Going down from Sonnet 5.5 is a cleaner case. Both models use the same tokenizer, so token counts carry over, and even Haiku 5.5's long tier ($0.50/$2.50) is a quarter of Sonnet 5.5's flat $2/$10. The short tier is a twentieth.

How Haiku 5.5 Scored Against GPT-6 Luna and Sonnet 5.5

On Anthropic's published evals, Haiku 5.5 sits well above Luna and below Sonnet 5.5 everywhere. The launch post reports, for Haiku 5.5 / GPT-6 Luna / Sonnet 5.5:

  • GDPval-AA v2.1 (Elo): 1620 / 1437 / 1840
  • OSWorld 2.1, offline subset: 72.4% / 48.9% / 83.9%
  • Terminal-Bench 4.0: 39.2% / 16.4% / 70.6%
  • FrontierCode 1.1 (Main): 46.4% / 42.4% / 52.1% (Sonnet at xhigh effort)
  • Chartography, no tools: 46.4% / 29.1% / 61.6%

Haiku 4.5 scored 15.7% on that OSWorld subset and 0.0% on Terminal-Bench 4.0, so this is a different class of model from the one it replaces. Two caveats. Luna is the only non-Anthropic model in the table, and OfficeChai's benchmark write-up notes Luna has no Humanity's Last Exam score in it, so you can't compare the two on that test. And Anthropic itself says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks."

Customer results in the launch post are vendor-selected but specific. HubSpot reported 92.8% averaged over three runs on its CRM suite. AlphaSense scored 0.84 against Haiku 4.5's 0.76 on 400 queries. Box saw 11 points over Haiku 4.5 at about half the latency. None of those were run against Luna. AlphaSense called its eight-point gap on 400 queries statistically significant; a two-point gap on a set that size usually isn't, the sample-size argument we made in Only 3 of 36 Model Gaps Were Real.


What Breaks When You Swap the Model ID

Haiku 5.5 isn't a drop-in replacement for Haiku 4.5, and several of the breaks land in classifier code. The migration guide and what's new page list them:

  • Any temperature other than 1, any top_k, or any top_p other than 0.99 returns a 400 error. A classifier pinned to temperature: 0 fails on the first call.
  • A final assistant turn (prefill, the old trick for forcing JSON) returns a 400 even with thinking off. Anthropic points you to structured outputs, or to tools with enum fields on Amazon Bedrock, which doesn't support structured outputs.
  • Adaptive thinking is on by default at medium effort. Responses can open with a thinking block, so code that reads the first content block as the answer breaks, and thinking tokens bill as output and count against max_tokens. A tight max_tokens tuned on Haiku 4.5 can stop before any text.
  • Safety classifiers can return stop_reason: "refusal", with no server-side fallback.
  • Priority Tier isn't supported. If you have a Priority Tier commitment on Haiku 4.5, plan capacity separately.

The thinking default is the one that changes cost. A 3,000-token classification prompt with a 200-token answer costs about $0.004 on Haiku 4.5. On Haiku 5.5, after 30% tokenizer inflation, it is about $0.0005. If adaptive thinking adds 1,000 output tokens to that call (an illustration; measure yours), the cost roughly doubles to about $0.001. That is still far below Haiku 4.5, but it is twice the no-thinking figure. Anthropic's guidance is to choose a lower effort level where Haiku 4.5 ran without thinking; you can also turn thinking off at high effort or below. We covered how effort settings move cost per task in Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score.

Two smaller items for the platform team. The model is on the Claude API, Bedrock, Google Vertex AI and Microsoft Foundry, with retirement "not sooner than October 7, 2027" per the model page. And US-only inference through inference_geo carries a 1.1x multiplier on Claude 4.6 and later models, per the pricing page, which Haiku 4.5 never supported.

Who Should Move, and Who Should Wait

Move short-prompt, high-volume traffic now and hold long-prompt work until you have run the numbers. The case for switching is strongest where Anthropic's 90% figure is true for you: intent routing, ticket triage, extraction from a single document, guardrail checks. Those prompts are short, the eval gains over Haiku 4.5 are large, and the per-text cost drops by close to an order of magnitude.

The strongest argument for waiting is that you may not need a frontier small model at all. For fixed-label classification, a decision model or fine-tuned classifier can undercut any of these per call, as our Jev vs Clef vs Haiku buyer's guide found. And if your workload is long-context retrieval or compaction, Luna's 272K line makes it cheaper per token on list price, before you test quality.

Teams running sub-agents on Sonnet 5.5 have the easiest decision: the tokenizer is the same, Haiku 5.5 costs a quarter as much even in its long tier, and Anthropic recommends it for that role. The risk is quality on the hardest steps, which is what a model router or a fallback rule is for.

This Week:

  1. Pull a week of Haiku 4.5 request logs and plot input tokens per request. Multiply by 1.3 and count how many land above 100,000. That share is your exposure to the 5x tier.
  2. Grep your codebase for claude-haiku-4-5 calls that set temperature, top_k or an assistant prefill. Each is a 400 error on Haiku 5.5.
  3. Send one request just over 100,000 tokens with a cached prefix and read the usage block to see which tier it billed.

This Month:

  1. Re-run your own eval set on Haiku 5.5 at low and medium effort, with cost per correct answer as the metric. Count tokens with the model set to claude-haiku-5-5, as the migration guide says, not from Haiku 4.5 counts.
  2. For compaction and long-context jobs, price the same workload on Haiku 5.5, Sonnet 5.5 and GPT-6 Luna, with each vendor's own tokenizer, before you move them.
  3. If you hold a Priority Tier commitment on Haiku 4.5, get a written answer from your account team on capacity for Haiku 5.5 before you migrate production.

The Bottom Line

Small-model pricing used to be a flat rate you could put in a spreadsheet. Haiku 5.5 brings length-tiered pricing, which OpenAI already applies at 272K tokens, down to the cheapest model in Anthropic's line, and the breakpoint sits at 100,000 tokens, where agent transcripts and retrieval prompts routinely cross. We saw the same pattern in OpenAI's GPT-6 Sol and Luna launch: the list price was halved and the real cost depended on settings the headline didn't show.

For routing, triage and extraction, Haiku 5.5 is a large and cheap upgrade over Haiku 4.5. For anything that reads a long transcript, the price you budget is $0.50 and $2.50. Get the prompt-length histogram before you change the model ID.

Continue Reading

Share:

Frequently Asked Questions

How much does Claude Haiku 5.5 cost?

Per Anthropic's pricing page, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Prompts over 100,000 tokens pay $0.50 input and $2.50 output, five times as much. Cache reads are $0.01 or $0.05, and the Batch API halves both tiers.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

Yes, but by less than the list price suggests. Haiku 4.5 is $1/$5 at any length. Haiku 5.5's tokenizer counts the same text as about 30% more tokens, so the same text costs roughly 87% less on short prompts and about 35% less on prompts over 100,000 tokens. Anthropic puts the average saving at about 75%.

Is Claude Haiku 5.5 better than GPT-6 Luna?

On Anthropic's own published evals it scores higher on every benchmark where both have a result, for example 72.4% vs 48.9% on an OSWorld 2.1 offline subset and 39.2% vs 16.4% on Terminal-Bench 4.0. Those are vendor numbers. List prices match below 100,000 tokens, but Luna stays at $0.10 input up to 272,000 tokens, so it is cheaper per token for long prompts.

What breaks when migrating from Claude Haiku 4.5 to Haiku 5.5?

Per Anthropic's migration guide, non-default temperature, top_p or top_k values return a 400 error, assistant prefill is rejected, manual budget_tokens thinking is replaced by adaptive thinking with an effort setting, responses can start with a thinking block, refusals can arrive with no fallback, and Priority Tier is not supported.

When should I keep using Claude Sonnet 5.5 instead of Haiku 5.5?

Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, and Sonnet 5.5 beats Haiku 5.5 on every benchmark in the launch table. Haiku 5.5 suits classification, routing, extraction, summarization and sub-agent work, where it costs a twentieth of Sonnet 5.5 on short prompts.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →