Anthropic Bid $7B for Cheaper Inference. Don't Fix Your Rate.

Anthropic is in advanced talks to buy Decart for roughly $7 billion, mostly in its own pre-IPO stock, to cut what a Claude token costs it to produce. Every Sonnet through 4.6 still lists at its March 2024 price, and the one cut that did land arrived as a new model number — which is why a flat multi-year rate card is the wrong thing to sign this quarter.

By Rajesh Beri·August 18, 2026·13 min read
Share:
A single printed contract page pinned flat under a heavy pen on a dark polished boardroom table, with a lit server rack visible through a glass wall in the background. No readable text or logos.

Illustration generated using AI

Your vendor just bid $7 billion to make the thing you are about to buy cheaper to produce. That should change what you sign, and it probably won't — because the clause that captures a falling cost curve is not in anyone's standard paper.

Anthropic is in advanced negotiations to acquire Decart, an Israeli AI-infrastructure company, for roughly $7 billion paid mostly in Anthropic shares, Calcalist reported on August 18. Advanced drafts have been exchanged and Calcalist says the deal could be signed as soon as next month. Weigh that against the earlier characterisation: when Bloomberg broke the story on August 13 at a $6 billion price, it reported the talks were at an early stage and could fall apart. Nothing is signed. If you are negotiating a multi-year Claude commitment this quarter, the deal is still not trivia. It is a dated, public signal about where your vendor's unit costs are heading — and a reason to refuse a flat rate card denominated in tokens.


What Anthropic Is Actually Buying for $7 Billion

Anthropic is buying a chip-efficiency layer and the ~100 people who built it, not a product line. Decart was registered on September 7, 2023 by Dean Leitersdorf and Moshe Shalev, has raised about $450 million, and was valued at roughly $4 billion three months ago, per Calcalist. At $7 billion that is a ~75% step-up on the May round, and about $70 million per employee.

The asset is DOS — the Decart Optimization Stack — which the company says "unlocks the full performance of every major chip" across NVIDIA, AWS Trainium and Google TPU hardware, according to Decart's own site. At the May round, Decart announced that DOS 2.0 runs agents at over 1,600 tokens per second — which its announcement put at eight times the industry average — and said it was generating significant revenue licensing DOS to cloud providers and AI labs. Treat both as vendor figures: they come from the company's own funding announcement, not an independent benchmark, and no third party has published a reproduction. Decart also ships consumer-facing world models — Lucy for real-time video transformation, Oasis for simulated environments — but Calcalist's reporting on the rationale puts the motive in cost terms: Anthropic, like its competitors, faces what it calls "the fundamental economic problem of frontier AI: enormous demand for computing comes with enormous costs," and extracting more work from the same silicon attacks that directly.

The most load-bearing detail is where the people land. Bloomberg's August 13 report — the one carrying the lower $6 billion figure and the early-stage caveat — said the Decart team would join Anthropic's inference and performance organization. Not research. Not product. The org whose entire job is making a token cost less to produce.

Why the Stock Consideration Is the Tell

Paying in pre-IPO shares tells you Anthropic thinks the cost reduction is worth more than the paper it is spending. Anthropic's annualized revenue run rate reached $65 billion at the end of July, up from $47 billion in May, against a $965 billion post-money valuation from its Series H — and the company has already filed confidentially for an IPO, with an internal projection of $190–200 billion in annual revenue by 2028.

Read those numbers together and the transaction is coherent. A company about to be marked by a public market has one number it must move: gross margin on inference. It has locked in enormous fixed capacity — Anthropic, Google and Broadcom announced multiple additional gigawatts of next-generation TPU capacity on April 6, 2026, coming online from 2027, on top of existing Trainium and TPU commitments. Fixed capacity plus a software layer that gets more tokens per chip is the fastest available path to margin. Buying it with stock, before the stock is publicly priced, is the cheapest way to pay for it.

Anthropic was not the only bidder. Calcalist reported that NVIDIA — already a Decart investor from the $300 million May round — put a higher valuation on the table and the process moved to Anthropic anyway, with Google and SpaceX also named as potential buyers. Anthropic, NVIDIA, Google and SpaceX have all been named around the same 100 people. That is the market telling you what inference efficiency is worth right now.

Anthropic's Price History Says Efficiency Reaches You Late

Cost reductions have historically reached Anthropic's customers as a new model number, not as a lower price on the model they already committed to. The evidence is on Anthropic's own pages.

Claude 3 Sonnet launched on March 4, 2024 at $3 per million input tokens and $15 per million output. Sonnet 4, Sonnet 4.5 and Claude Sonnet 4.6 are all still $3/$15 today, per Anthropic's pricing documentation. That is nearly two and a half years of flat list pricing across four model generations, through the largest efficiency gains in the industry's history. The streak broke only with Sonnet 5 in June 2026 — and it broke the way this whole section describes: on a new model number.

The Opus tier tells the other half. Claude 3 Opus was $15/$75 in March 2024; Claude Opus 4.5, announced November 24, 2025, came in at $5/$25 — a threefold cut. Anthropic framed the gain in tokens, not dollars: at its highest effort level, Opus 4.5 "exceeds Sonnet 4.5 performance by 4.3 percentage points—while using 48% fewer tokens."

Steel-man the vendor here, because there is a real and recent case. Claude Sonnet 5 launched on June 30, 2026 at $2/$10 — a 33% cut on the per-token rate Sonnet had held since 2024 — and Anthropic's docs record that the increase to $3/$15 scheduled for September 1, 2026 will not occur: the introductory rate became the standard rate. Prompt caching reads cost 10% of base input, the Batch API is a flat 50% off, and the full 1M-token context window carries no premium. Anthropic does pass savings through. It just does it on its own schedule, at the moment it ships a new model, to whoever is paying list — not to the customer who fixed a rate in a two-year agreement.

The Tokenizer Moved and Your Rate Card Didn't

A price per token is not a price, because the vendor defines the token. This is the part most procurement teams have not priced, and it is sitting in public documentation.

Anthropic's pricing page carries a note that Claude 4.7 and later models "use a newer tokenizer" that "produces approximately 30% more tokens for the same text," with the exact increase depending on content and workload shape. Sonnet 4.6 and earlier use the previous tokenizer. Now do the arithmetic on the vendor's own published figures: moving from Sonnet 4.6 at $3/$15 to Sonnet 5 at $2/$10 is a 33% cut per token, but the same body of text now bills roughly 30% more tokens. Multiply the two and the headline 33% price cut lands as roughly 13% on the same body of text. Text is not the same thing as work — a stronger model may finish a task in fewer passes — which is precisely why the corpus below, not the token, is the unit worth contracting on.

That is not an accusation of bad faith — a better tokenizer is a legitimate engineering choice, and Anthropic documented it. It is a demonstration that the unit is not stable, and a multi-year commitment denominated in dollars-per-million-tokens is a commitment to a unit your counterparty controls the definition of. The same trap runs the other way through effort settings, reasoning budgets, and per-model tool-use overhead: the pricing docs list a tool-use system prompt costing 286 tokens on Opus 5 and 675 on Opus 4.7 for the identical request. We have made this argument before about comparing frontier models on per-token price; the Decart deal is the reason it now belongs in the contract rather than the evaluation memo.

Two more multipliers stack on top and belong in any modelled rate: US-only inference via inference_geo carries a 1.1x multiplier on every token category, and regional or multi-region endpoints on Amazon Bedrock and Google Vertex AI add a 10% premium over global. If your data-residency posture is non-negotiable, your effective rate is 10% above the number on the page before you negotiate anything — the same structural surprise buyers hit with Mistral's regional European endpoints.

The Clauses That Survive a Price Collapse

Ask for a downward price-adjustment mechanism and a stable unit of work, in that order. Everything else is secondary.

Start from what the published paper actually says, because it is the floor you negotiate up from. Anthropic's Commercial Terms of Service provide that "Anthropic may update the published rates, to be effective the earlier of 30 days after the updates are posted by Anthropic or Customer otherwise receives Notice" (§H.1), and claude.com/pricing states plainly that price and plans are "subject to change at Anthropic's discretion." The same terms let Anthropic assign the agreement "to an affiliate or as part of a sale of all or substantially all its business" without your consent (§M.4). A standard change-of-control clause does nothing for you here — the vendor being acquired in this story is Decart, not Anthropic.

The four asks that matter:

  1. A ratchet, not a rate. Your negotiated discount applies to list, and if list falls, your price falls with it. Never the reverse. A fixed dollar rate across a 24- or 36-month term is a bet against your vendor's own proposed $7 billion capital allocation.
  2. Define the unit as work, not tokens. Pin a benchmark corpus — a few hundred of your real prompts and expected outputs — into the agreement, and define a price change as a change in cost to process that corpus. A tokenizer swap, a default effort-level change, or a mandatory reasoning budget that moves your corpus cost by more than a stated threshold is a price change and triggers the same 30-day notice and the same adjustment right.
  3. Model-substitution at successor pricing. You commit spend, not a SKU. When the successor model ships, committed dollars move to it at its list price, without renegotiation and without your discount resetting. Otherwise the next Sonnet-5-shaped cut lands on everyone except you.
  4. Drawdown, not prepay-and-forfeit. Unused committed spend rolls into the next period or converts. If efficiency gains mean you consume half the tokens you forecast, a use-it-or-lose-it commit converts your vendor's cost win into your write-off — exactly the repricing exposure buyers hit when Bending Spoons acquired Airtable and inference credits were re-cut mid-term.

Anthropic's own pricing page confirms this is negotiable territory: the sales-assisted Enterprise plan explicitly supports "MSA, PO, usage commitments, product bundling," and volume discounts are described in the pricing docs as negotiated case by case. Self-serve Enterprise is $20 per seat plus usage at API rates. If you are on the self-serve tier, you have no adjustment mechanism at all — you have a published price list and 30 days' notice.

If You License DOS, You Have a Different Problem

Decart's existing customers face a change-of-control question, and it is sharper than the usual one because of who those customers are. SiliconANGLE's reporting says DOS revenue comes from licensing to cloud providers and AI labs. If that is you, your chip-efficiency vendor may be about to be owned by a frontier lab you compete with, and the "hardware-agnostic" property you bought — DOS spanning NVIDIA, Trainium and TPU — is now a strategic decision made by that lab rather than a product commitment.

We wrote the general version of this when d-Matrix acquired Wallaroo: get "any hardware" in writing, with named platforms and a support horizon, before the deal signs. The same applies to Decart's API customers. Lucy and Oasis are sold pay-as-you-go today, priced per second by model — $0.02/sec realtime and $0.04/sec of generated 720p video for Lucy 2.5, $0.01/sec for Lucy Restyle 2, with no subscriptions and no minimum spend. "No minimum spend" is a nice consumer term and a terrible enterprise one — it means nothing constrains the acquirer's pricing after close. Anthropic has until now largely stayed out of Israel physically but is laying the groundwork for a larger push there, having acquired the AI-infrastructure startup Runhouse four months ago, so treat continuity of the standalone products as unresolved rather than assumed.

NVIDIA's position is the detail to watch. It is simultaneously a Decart investor, a losing bidder, and the vendor of most of the silicon DOS optimises. Ask your Decart account team, in writing, for a named support horizon on non-NVIDIA and non-Anthropic-preferred hardware paths.

What to Do Before You Sign

This Week:

  1. Pull every AI vendor agreement with a term longer than 12 months and grep for a price-adjustment clause. Most have an escalator and no ratchet. Write down which ones expire before Anthropic's expected listing window.
  2. Freeze a benchmark corpus — 200 to 500 real production prompts with expected outputs — and run it against your current model. Record total input tokens, output tokens and dollar cost. That single number is your unit of account for every negotiation from here.
  3. If a Claude commitment is mid-negotiation, tell your account team the signing target for the Decart transaction is next month and ask what happens to your rate if inference costs fall. Get the answer in the redline, not on the call.

This Month:

  1. Re-run the benchmark corpus on the newest model tier and compare cost per completed task, not cost per token. If the token count moved more than the price did, you have found your clause.
  2. Instrument per-model, per-task cost in your gateway so the corpus number is continuously measured rather than a one-off — the gateway comparison covers what to look for, and prompt caching hit rates matter more than routing here, as we found when caching broke the multi-model router economics.
  3. Inventory whether anything in your stack depends on Decart — directly, or through a cloud provider or model vendor that licenses DOS. Vendor-of-your-vendor exposure is the disclosure gap that keeps recurring, most recently with Stripe's acquisition of OpenRouter.

Before Renewal:

  1. Refuse any fixed per-token rate longer than 12 months without a downward adjustment mechanism. Take a shorter term at a worse headline discount over a long term at a frozen rate. In a market where the vendor just bid $7 billion of its own equity on cost reduction, optionality is worth more than three points of discount.

The Bottom Line

This is the reserved-instance trade, replayed with a worse-defined unit. Cloud buyers learned in the 2010s that a three-year commitment at today's rate is a bet the provider's costs will stop falling — and the providers kept shipping cheaper instance families to everyone except the people who had prepaid. The AI version is harder, because at least a vCPU-hour meant the same thing in year three as it did in year one. A token does not: Anthropic's own documentation says its newer tokenizer produces about 30% more of them for the same text.

The company that just committed multiple gigawatts of TPU capacity from 2027 and is negotiating to pay roughly $7 billion in stock for a chip-efficiency team is telling you, with its capital, that it expects to produce a token far more cheaply than it does today. Believe it. Then make sure your contract is written in a unit that lets you collect.

Price the work. Never the token.

Continue Reading

Share:

Frequently Asked Questions

Why is Anthropic acquiring Decart?

To lower its inference cost base. Decart's DOS (Decart Optimization Stack) extracts more throughput from NVIDIA, AWS Trainium and Google TPU hardware, and Bloomberg reported the team would join Anthropic's inference and performance organization. Bloomberg put the price at $6 billion on August 13 and reported the talks were at an early stage and could fall apart; Calcalist reported roughly $7 billion on August 18, mostly in Anthropic shares, with advanced drafts exchanged and signing possible as soon as September 2026. The deal is not signed.

Should I sign a multi-year Claude commitment right now?

Only with a downward price-adjustment mechanism. A vendor spending $7 billion of its own equity to cut production costs is signalling that its prices can fall, and a fixed per-token rate across 24 to 36 months captures none of that. Take a shorter term at a worse headline discount over a long term at a frozen rate.

Has Anthropic ever cut its API prices?

Yes, but at the model tier, not on existing commitments. Claude 3 Opus was $15/$75 per million tokens in March 2024; Opus 4.5 launched at $5/$25 in November 2025. Sonnet held $3/$15 across four generations from March 2024 — Sonnet 4, 4.5 and 4.6 are all still sold at that price — until Sonnet 5 launched on June 30, 2026 at $2/$10. Anthropic later cancelled the scheduled September 1, 2026 increase, making $2/$10 permanent. Every one of those cuts arrived as a new model number, not as a reduction on a model a customer had already committed to.

Why is a price per token not a real price?

Because the vendor defines the token. Anthropic's pricing documentation states that Claude 4.7 and later models use a newer tokenizer producing approximately 30% more tokens for the same text. On those published figures, the 33% per-token cut from Sonnet 4.6 to Sonnet 5 lands as roughly 13% on the same body of text. Denominate commitments in cost to process a fixed benchmark corpus instead.

What should a Decart customer do before the deal signs?

Get hardware coverage and a support horizon in writing. DOS is licensed to cloud providers and AI labs and spans NVIDIA, Trainium and TPU today; after close, that portfolio becomes a strategic decision made by a frontier lab. Decart's Lucy and Oasis APIs are sold pay-as-you-go with no minimum spend, so nothing currently constrains post-close repricing.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →