Inference Cost per Million Tokens: Price the Task, Not the Rate

Priced on one normalised agent workload, a month of inference costs $370 on GPT-5.6 Luna and $10,725 on Claude Opus 5. Output, reasoning and cache-read rates explain the gap, not the input price.

By Rajesh Beri·September 11, 2026·22 min read
Share:
A long curling paper receipt spilling out of a desktop adding machine across a wooden office desk, a short stretch near the top highlighted in yellow and a much longer stretch further down highlighted in orange, with a p

Illustration generated using AI

On our reference agent workload, output and reasoning tokens are roughly 60% of the bill at almost every major provider, so the input price everyone quotes is the least useful forecasting number. Priced on one identical workload, a month of inference runs from $370 on GPT-5.6 Luna to $10,725 on Claude Opus 5, and inside the frontier tier alone the spread is 3.6x. The rate card is where a forecast starts. It is not where the money goes.

Every price below was checked on 11 September 2026 against each vendor's live pricing page. Two of those prices already carry a printed expiry date, and a third vendor retires a model on 14 September. Treat this page as a snapshot with a date on it, because that is what every inference price is.

The Verdict: Rank on Output, Cache Reads and Discounts

Forecast on output price, cache-read price and discount eligibility, in that order — and on those three, here is how the major providers land:

  • The loser on cost is Claude Opus 5. It has the highest rate in its tier, a tokenizer that bills roughly 30% more tokens for the same text, and on the same benchmark it generated 56% more output tokens than GPT-5.6 Sol. Buy it only where your own eval shows it wins.
  • Cheapest frontier rate for interactive work: Grok 4.6, at $2,950 a month on our workload. It is also the only frontier model whose price does not move when you batch, and one token past 200,000 doubles the whole request.
  • Cheapest frontier batch lane: Gemini Pro — specifically Gemini 3.1 Pro Preview, at $2,000 a month batched. That is a Preview label on a production budget, and you should price it as one.
  • Volume: Gemini Flash — Gemini 3.8 Flash — for the rest of 2026, budgeted at double from 1 January 2027. Google prints the step on its own pricing page. After it lands, Grok 4.3 is cheaper on this workload.
  • Floor: GPT-5.6 Luna, unless DeepSeek's off-peak rate clears your data-residency review.
  • Claude Sonnet 5 vs GPT-5.6 Terra: Terra, unless your measured tokenizer ratio comes in under 1.12x.

The workload, stated so you can argue with it: an internal agent running 100,000 tasks a month, each reading a 20,000-token context of which 15,000 tokens are a shared, cached prefix, and producing 2,000 output tokens including reasoning. That is 1,500 million cached input tokens, 500 million fresh input tokens and 200 million output tokens a month, counted on OpenAI's tokenizer. Rows for Claude 4.7 and later are then multiplied by 1.30, Anthropic's own tokenizer disclosure (explained below). That ratio compares Anthropic's new tokenizer with its previous one, not with OpenAI's, so the table assumes the older Claude tokenizer and every other vendor's count roughly the same as OpenAI's. At this prompt size, OpenRouter's billing data puts the effective Claude increase at 21–25%, so read the Claude rows as the high end of a range. We assume the prefix stays warm, so cache-write premiums are excluded, and no negotiated discount is applied.

Tier Model Input / cached / output, per MTok Batch Reference workload, per month Verdict
Frontier Claude Opus 5 $5 / $0.50 / $25 −50% $10,725 ($8,250 before tokenizer) Loser on cost; escalation only
Frontier GPT-5.6 Sol $4 / $0.40 / $20 (promotional) −50% $6,600 ($9,250 at $5/$30) Budget both rates
Frontier Gemini 3.1 Pro Preview $2 / $0.20 / $12 (prompts ≤200K) −50% $3,700 ($2,000 batched) Cheapest frontier batch lane
Frontier Grok 4.6 $2 / $0.50 / $6 (prompts <200K) none $2,950 Cheapest interactive; no batch
Mid Claude Sonnet 5 $2 / $0.20 / $10 −50% $4,290 ($3,300 before tokenizer) Loses to Terra above 1.12x
Mid GPT-5.6 Terra $2 / $0.20 / $12 −50% $3,700 Safest second provider
Mid Mistral Medium 3.5 $1.50 / $0.15* / $7.50 −50% $2,475 Flattered by the cache assumption
Mid Grok 4.3 $1.25 / $0.20 / $2.50 −20% $1,425 Cheapest mid-tier from January
Mid Gemini 3.8 Flash $0.75 / $0.075 / $3.75 −50% $1,238 → $2,475 from 1 Jan 2027 Volume default for 2026
Volume Claude Haiku 4.5 $1 / $0.10 / $5 −50% $1,650 4.5x the floor
Volume Mistral Large 3 $0.50 / $0.05* / $1.50 −50% $625 Flagship name, small-model price
Volume Gemini 3.1 Flash-Lite $0.25 / $0.025 / $1.50 −50% $463 Google's floor
Volume DeepSeek V4.1 Flash $0.30 / $0.006 / $1.20 (peak) off-peak −50% $399 peak, $200 off-peak Only if residency allows
Volume GPT-5.6 Luna $0.20 / $0.02 / $1.20 −50% $370 The floor

Rate cards from Anthropic's pricing documentation, OpenAI's API pricing, the Gemini API pricing page, xAI's model pricing, DeepSeek's pricing page and Mistral's API pricing, all read 11 September 2026. *Mistral says cached input tokens cut "input costs by 90%" but does not say which models qualify; I applied it at face value, which flatters both Mistral rows. Tiers follow each vendor's own flagship/mid/small positioning. This page prices models; it does not rank their quality, and the cheapest model that fails your eval is not cheap.


Why Output and Reasoning Tokens Decide the Bill

Output tokens are 9% of the tokens in this workload and about 61% of the bill on Opus 5, Sol, Sonnet 5 and Gemini 3.8 Flash — 65% on Terra and Luna. Cached reads are the mirror image: 68% of the tokens, 8–9% of the bill wherever the cache-read rate is 10% of input. Across all 14 rows, output lands between 60% and 65% of the bill on 11 of them. The three exceptions are the two Grok models, whose output is unusually cheap, and Mistral Large 3. That share is a property of this workload, not a law: a coding agent that re-reads a 150,000-token cached transcript to write 1,000 tokens flips it, and on Opus 5 cache reads become three-quarters of the bill.

A reasoning token is a token the model generates to think before it answers — you never see it, and you pay for it at the output rate. OpenAI's documentation says it plainly: reasoning tokens "are not visible via the API" but "still occupy space in the model's context window and are billed as output tokens," and it recommends reserving at least 25,000 tokens for reasoning and output when you start. Google prints its output rates as "including thinking tokens". Anthropic exposes a thinking_tokens field that reports "how many of the billed output tokens were internal reasoning".

That makes reasoning volume a pricing variable as large as the rate itself, and it varies by model on identical work. Artificial Analysis reports that Claude Opus 5 at max effort used 140 million output tokens to run its Intelligence Index, and cost $7,274.74 to evaluate. GPT-5.6 Sol at max effort used 90 million and cost $3,464.84. Opus 5's output rate is 1.25x Sol's. Its bill for the same test was 2.1x. That gap is the rate multiplied by verbosity, counted in each vendor's own tokens, so Anthropic's tokenizer is part of it, and no rate card shows it.

The sensitivity is worth putting in the forecast as its own row. At 100,000 tasks a month, every additional 1,000 reasoning tokens per task adds 100 million output tokens: $2,000 a month on Sol, $2,500 on Opus 5, $600 on Grok 4.6, $375 on Gemini 3.8 Flash and $120 on Luna. On Sol that is as much as the entire fresh-input line. Raising the effort setting one notch on a frontier model can move the bill more than switching vendors would.

Two Anthropic mechanics add hidden input on top. Claude Opus 4.5 and models numbered 4.6 and later keep earlier turns' thinking blocks in context and bill them as input, so an agent's reasoning is paid for once as output and again on every later turn. And changing the thinking configuration between requests invalidates the prompt cache, so a router that flips effort levels per request pays full input price on the context it just broke.

Cached Input Is Where the Rate Cards Really Differ

On a workload where two-thirds of the tokens are cache reads, the cache-read multiplier matters more than the headline input price, and it ranges from 2% to 25% of base input across these vendors. Prompt caching is the provider storing a processed prefix — system prompt, tool definitions, documents, conversation so far — and re-serving it at a discount on the next request.

The read multipliers, as of 11 September 2026:

  • DeepSeek V4.1 Flash: 2% — $0.006 against $0.30 at peak.
  • Claude Fable 5.1 and Mythos 5.1: 2.5% — every other Claude model uses 10%.
  • OpenAI, Anthropic (other models) and Google: 10%.
  • Grok 4.3: 16% ($0.20 against $1.25).
  • Grok 4.6: 25% ($0.50 against $2.00). This is why Grok 4.6's cache line is a quarter of its bill on our workload, against 8–9% elsewhere.

The write side differs too. Anthropic charges 1.25x base input for a five-minute cache write and 2x for a one-hour write. OpenAI's caching is automatic, but for GPT-5.6 and later "cache writes cost 1.25× the standard, uncached input-token rate", a prefix stays eligible for 30 minutes after its last use, and the minimum cacheable prompt is 1,024 tokens. Google's implicit caching is on by default for Gemini 2.5 and newer, but only above a minimum of 4,096 tokens on the Gemini 3.x Flash models, and explicit caching adds a storage charge per million tokens per hour — $4.50 on Gemini 3.1 Pro Preview — whether or not you send a request.

The premium tier shows how much this matters. Claude Fable 5.1 and GPT-6 Astra carry the same $10 input and $50 output rate card. Fable 5.1 reads cache at $0.25; Astra at $1.00. On our workload Astra costs $16,500 a month and Fable 5.1 $15,375 at face value — and $19,988 once Anthropic's tokenizer is applied. The cache discount is real; on this mix it does not survive the token count. We worked through where it does in the Fable 5.1 cache-read analysis.

Batch, Flex and Off-Peak: The 50% Nobody Budgets

Any request that can wait up to 24 hours costs half as much at OpenAI, Anthropic, Google, Mistral and Amazon Bedrock — and almost no forecast assumes it. Batch is an asynchronous queue: you submit requests, and results come back within a window, usually a day, at a discount.

  • OpenAI prices Batch and Flex at half of Standard across the GPT-5.6 family — Sol drops to $2 / $0.20 / $10.
  • Anthropic gives a 50% discount on input and output, and states that its caching multipliers "stack with other pricing modifiers, including the Batch API discount."
  • Google halves input and output prices for Batch and Flex on the Gemini 3.x models.
  • Mistral offers batch "at half price."
  • Bedrock offers batch inference "at a 50% lower price" on select models, and a Flex tier at a 50% discount.
  • xAI is the outlier. Its batch discount is 20% and applies to four models — Grok 4.3 and three Grok 4.20 variants — and models not on that list get no batch discount. Grok 4.6 is not on it.

Batching reorders the frontier tier. On our workload, Gemini 3.1 Pro Preview falls to $2,000, Sol to $3,300 and Opus 5 to $5,363 after the tokenizer adjustment. Grok 4.6 stays at $2,950. It goes from the cheapest interactive frontier model to 48% more expensive than Gemini 3.1 Pro Preview once the work can wait.

DeepSeek does the same thing with a clock instead of a queue. Its off-peak rates are half of peak, and peak is 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday — 35 of the week's 168 hours. For a US company that is the local evening and overnight, so US business-hours traffic pays off-peak prices all day without a line of code changing. European mornings fall squarely in peak.

Speed costs the other way. OpenAI's fast mode doubles Sol to $8 input and $40 output. Anthropic's fast mode puts Opus 5 at $10 input and $50 output. Google's Priority tier is 1.8x Standard — $1.35 input on Gemini 3.8 Flash. Bedrock's Priority tier is a 75% premium. Every latency guarantee is a second rate card, and it belongs in the forecast if anyone in the organisation can switch it on.

Where Context Length Doubles the Price

Two of these vendors reprice the entire request once a prompt crosses 200,000 tokens, and Anthropic charges the same per-token rate all the way to a million. Context length is the number of input tokens in one request, and on agents it grows every turn.

xAI's pricing rule is that "requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request", and for Grok 4.6 that threshold is 200,000 tokens, where the rate goes from $2 / $6 to $4 / $12. Google's Gemini 3.1 Pro Preview moves from $2 / $12 to $4 / $18 for prompts over 200,000 tokens. Anthropic states that Claude 4.6 and later models include the full 1M-token context window at standard pricing: "A 900k-token request is billed at the same per-token rate as a 9k-token request." OpenAI prices GPT-5.6 Sol prompts above 272,000 input tokens at "2x input and 1.5x output for the full request" — $8 / $30 — and lists GPT-6 Astra with a separate long-context rate of $20 input and $75 output.

The cliff is literal. One uncached request with 2,000 output tokens:

Prompt size Grok 4.6 Gemini 3.1 Pro Preview Claude Opus 5
199,000 tokens $0.410 $0.422 $1.045
201,000 tokens $0.828 $0.840 $1.055

One percent more context doubles the Grok and Gemini requests and moves Opus 5 by one percent. If your agents run long, put a hard guard at 195,000 tokens on those two models, or route anything past it elsewhere.

Context also compounds through the cache line. An agent that re-reads its whole transcript every turn pays for every earlier turn again, at the cache-read rate if it is lucky and the full rate if the cache went cold. That is how one pull request billed 156 million tokens. Tool definitions ride along too: Anthropic documents that its computer-use toolset adds about 4,500 input tokens to any request that declares it, and its browser toolset about 6,600.

How to Compare Prices When Tokenizers Differ

A per-token price comparison is only valid if both vendors turn your text into the same number of tokens, and they do not — so compare cost per task, measured on your own requests. A tokenizer is the function that splits text into the billable units a model charges for, and each vendor's is different.

Anthropic discloses the biggest known gap in a note on its pricing page: Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text," with the exact increase depending on content and workload shape. Claude Sonnet 4.6 and earlier — which includes Haiku 4.5 — use the previous tokenizer. So even inside one vendor, comparing Haiku 4.5 with Sonnet 5 on the rate card understates the step.

The effective impact depends on prompt length. OpenRouter analysed more than a million requests across the change and found 32–45% raw inflation, but an effective billing change of −1.6% under 2,000 tokens, +27.2% at 2,000–10,000, +21–25% at 10,000–50,000, and +12–15% above 50,000, where caching absorbs most of the extra tokens. Content type matters as much. One measurement through Anthropic's token counter found 1.47x on English technical documentation but 1.07x on numeric CSV and 1.01x on Japanese and Chinese prose, and Simon Willison measured 3.01x on a high-resolution image.

Here is what that does to the table. Sonnet 5 at face value is $3,300 against Terra's $3,700 — 11% cheaper. At 1.30x it is $4,290 and 16% more expensive. The break-even on this workload is 1.12x: below it Sonnet 5 wins, above it Terra does. Opus 5 needs no adjustment to lose to Sol — it is 25% more expensive at face value and 62.5% more after the tokenizer.

No other vendor on this list publishes a comparable ratio, which is not the same as their tokenizers matching OpenAI's. The only honest number is the one you measure: run 500 real production requests through Anthropic's token-counting endpoint, OpenAI's usage counts and Google's countTokens, per content type, and divide. Then forecast in cost per task or per 1,000 characters, never in cost per token.

Committed-Use Discounts: Commit Dollars, Not Rates

A committed-use discount only saves money if list prices do not fall faster than the discount, and in 2026 they have repeatedly fallen faster. A committed-use discount is a lower rate in exchange for a promise to spend, or to hold capacity, for a fixed term.

The first-party APIs do not publish theirs. Anthropic says volume discounts "are negotiated on a case-by-case basis"; on AWS Marketplace, a negotiated discount is applied when token usage converts to Claude Consumption Units. The clouds publish structures:

Capacity commitments carry a utilisation trap. Provisioned capacity bills every hour whether or not a token passes through it, so its real unit cost is the hourly rate divided by the tokens you actually push. At 50% utilisation, a 50% term discount buys you exactly the no-commitment price.

Rate commitments carry a price-decline trap. Anthropic launched Sonnet 5 at $2 / $10 against Sonnet 4.6's $3 / $15. OpenAI cut Sol's input 20% and output 33.3%, to $4 / $20 from $5 / $30, on a promotion that runs "at least through November 21, 2026." Anthropic cut Fable 5.1's cache reads to a quarter of Fable 5's. A 12-month commitment at 20% off Sol's earlier $30 output rate would now pay $24 per million output tokens against a $20 list price. We made the same argument about fixing a rate while a vendor's own inference costs are falling.

Residency is a commitment too. US-only inference costs 1.1x on Claude 4.6 and later, and regional Claude endpoints on Bedrock and Google Cloud carry a 10% premium. OpenAI charges a 10% uplift for regional processing on models released on or after 5 March 2026. Mistral's Enterprise APIs, which include regional data processing controls, are "75% above list pricing on select APIs" — which erases most of Mistral's price advantage for the EU buyers most likely to want it, a pattern we saw in its European compute commitments.

Who Should Not Buy Each Option

Every model on this list is the wrong answer for someone, and the rate card never says who:

  • Claude Opus 5 — anyone routing volume to it, and anyone whose eval does not show a measurable win over Sonnet 5 or Terra on the specific task.
  • GPT-5.6 Sol — anyone sizing an FY27 commitment on $4 / $20 without a second column at $5 / $30, which puts our workload at $9,250 instead of $6,600. We covered the reversion arithmetic in the Sol promotion analysis.
  • Gemini 3.1 Pro Preview — anyone with prompts that cross 200,000 tokens, anyone holding a large explicit cache warm, and anyone whose procurement policy excludes Preview models from production.
  • Grok 4.6 — batch-heavy pipelines, long-context agents, and cache-heavy loops, where its 25% cache-read rate costs 2.5x the 10% norm.
  • Claude Sonnet 5 — English documentation and prose workloads, where the tokenizer penalty runs well above 1.12x.
  • GPT-5.6 Terra — any workload where Gemini 3.8 Flash passes the eval; Terra costs 3x as much this year and 1.5x after Google's price step.
  • Gemini 3.8 Flash — anyone forecasting into 2027 at the current rate, and bursty traffic with prompts under 4,096 tokens, which never hits the implicit cache.
  • Mistral Medium 3.5 and Large 3 — anyone who needs Mistral's regional data processing controls at the 75% enterprise uplift, and anyone who cannot confirm the cache discount applies to their model.
  • DeepSeek V4.1 Flash — regulated data, anyone with a China-jurisdiction restriction, and European-morning traffic that pays peak. We laid out the supply-chain and residency questions separately. Note also that DeepSeek is retiring deepseek-v4-pro and rerouting its requests to V4.1 Flash from 14 September.
  • GPT-5.6 Luna and Claude Haiku 4.5 — tasks that need multi-step reasoning. And Haiku 4.5 specifically for pure classification or extraction, where Luna costs less than a quarter as much on our workload.

The Decision Criteria That Predict Regret

Six properties of your workload decide which vendor is cheapest, and all six can be measured from a week of logs before you sign anything:

  1. Output share of the bill. If output plus reasoning is above half your spend, rank on output price and nothing else first.
  2. Cache hit rate. Above about 60% of input tokens from cache, the cache-read multiplier outranks the input price. This is where Grok 4.6 loses its advantage and DeepSeek and Fable 5.1 gain.
  3. Reasoning volume per task at your effort setting. Measure it from reasoning_tokens or thinking_tokens; do not take it from a benchmark.
  4. Prompt length distribution. If any meaningful share of requests crosses 200,000 tokens, Grok and Gemini Pro pricing doubles for those requests, and Claude's flat long-context rate starts to win.
  5. Latency tolerance. The share of traffic that can wait a day is the share that costs half — everywhere except Grok 4.6.
  6. Residency requirement. It adds 10% at OpenAI, Anthropic and the cloud platforms, and up to 75% at Mistral.

What changes the answer: if more than half your traffic is batch-eligible, Gemini 3.1 Pro Preview beats Grok 4.6 in the frontier tier. If your measured Claude tokenizer ratio is under 1.12x, Sonnet 5 beats Terra. If your forecast runs past 31 December 2026, Grok 4.3 undercuts Gemini 3.8 Flash on this workload. And if an AI gateway is already in your stack, routing by these criteria is a configuration change rather than a migration — see our gateway comparison, and mind the cache you forfeit when a router switches models.

What to Do Before the FY27 Budget Locks

This Week:

  1. Split last month's bill into four lines per model — cached input, fresh input, visible output and reasoning output — using cached_tokens, reasoning_tokens and thinking_tokens from the usage objects. If you cannot produce the fourth line, you do not know what you are buying.
  2. Run 500 real requests through each candidate vendor's token counter and record the ratio per content type. Replace the 1.30x in this article with your number.

This Month:

  1. Tag every workload that can tolerate a 24-hour turnaround and move it to batch or flex. Start with evals, backfills and nightly enrichment.
  2. Put a 195,000-token guard on any Grok 4.6 or Gemini 3.1 Pro Preview route.
  3. Price one effort-level change on your most expensive agent before anyone raises it in production.

Before Budget Lock:

  1. Add two dated rows to the forecast: Sol at $5 / $30 after 21 November 2026, and Gemini 3.8 Flash at double from 1 January 2027. Both are printed on the vendors' own pages; neither is a surprise.
  2. Commit dollars, not rates: a spend drawdown usable across a vendor's model family, a term of 12 months or less, and a clause that passes list-price cuts through to your committed rate.

The Bottom Line

The early cloud market taught this lesson once already. The instance-hour price was on every slide, and the bill was set by egress, storage and idle capacity nobody had modelled. Inference is repeating it with tokens: the input price is on every comparison chart, and the bill is set by output, reasoning, cache reads, context cliffs and a tokenizer nobody measured.

A per-million-token price is a unit of account, not a forecast. The providers on this page differ by 29x on the same month of work, and most of that gap lives in lines the rate card prints in small type or not at all.

The rate card tells you what a token costs. Your workload decides how many you buy. Forecast the second number.

Continue Reading

Share:

Frequently Asked Questions

What is the cheapest LLM API per million tokens in 2026?

As of 11 September 2026, GPT-5.6 Luna ($0.20 input / $1.20 output per million tokens) and DeepSeek V4.1 Flash ($0.30 / $1.20 at peak, half that off-peak) are the lowest-priced models from major providers. On a 100,000-task agent workload, Luna costs about $370 a month and DeepSeek about $200 off-peak. The cheapest rate is not always the cheapest task once output and reasoning volume are counted.

Are reasoning tokens billed as output tokens?

Yes. OpenAI states reasoning tokens are not visible via the API but are billed as output tokens, Google prices Gemini output 'including thinking tokens', and Anthropic reports billed reasoning in a thinking_tokens usage field. Because they bill at the output rate, reasoning volume can move a bill more than switching vendors.

How do I compare LLM prices when tokenizers differ?

Measure your own text. Anthropic says Claude 4.7 and later produce about 30% more tokens for the same text, so a per-token comparison is not like-for-like. Run a few hundred real requests through each vendor's token counter, record the ratio per content type, and compare cost per task or per 1,000 characters instead of cost per token.

Which LLM providers offer a batch API discount?

OpenAI, Anthropic, Google, Mistral and Amazon Bedrock all offer 50% off for asynchronous batch processing. xAI offers 20% on Grok 4.3 and three Grok 4.20 variants and no batch discount on Grok 4.6. DeepSeek instead halves its prices during off-peak hours, which cover all of US business hours.

Does a longer context window cost more per token?

It depends on the vendor. Claude 4.6 and later models bill the full 1M-token window at the standard rate. Gemini 3.1 Pro Preview and Grok 4.6 switch to a higher rate once a prompt crosses 200,000 tokens, roughly doubling the cost of that request.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →