If you turned on prompt caching and your bill barely moved, the cache is almost certainly missing, and something in your own prompt is the likely cause. On a 30-turn agent loop priced from each vendor's live page, a correctly cached run saves 80% on Claude Sonnet 5.5, 85% on GPT-6.1 Sol and 89% on DeepSeek V4 Pro. Put a timestamp in the system prompt and the Claude saving falls to 6%. Put it above the tool definitions and you pay 24% more than with caching switched off, because cache writes cost more than plain input.
The verdict: turn caching on everywhere, then spend your engineering time on prefix stability, not on picking a provider. Among the frontier APIs, GPT-6.1 Sol has the cheapest cache read. Claude gives you the most control and keeps hits off your rate limit. DeepSeek comes closest to the advertised 90% because it charges nothing to write. Gemini 3.1 Pro is the loser here: it has the highest minimum prefix and an implicit-hit discount its caching page does not quantify, and the one independent study of agent workloads measured its predecessor, Gemini 2.5 Pro, saving roughly half what the others did.
All prices below were checked on each vendor's live pricing page on October 2, 2026.
| Model (standard tier) | Cache read /MTok | Cache write /MTok | Min. cacheable prefix | Cache lifetime | Our 30-turn agent loop: uncached → cached | Who should NOT pick it for caching |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | $0.20 | $2.50 (5 min), $4.00 (1 h) | 512 tokens | 5 min, or 1 h at 2x | $2.72 → $0.53 (-80%) | Teams whose agent turns routinely outlast 5 minutes and who won't pay for the 1-hour write |
| GPT-6.1 Sol | $0.10 | $2.50 | 1,024 tokens | 30 min | $2.72 → $0.41 (-85%) | Anyone who has to promise finance a hit rate, or pushes one prefix past ~15 requests a minute |
| Gemini 3.1 Pro Preview | $0.20 (≤200K) | none on implicit; explicit caches bill $4.50/MTok-hour storage | 4,096 tokens | Implicit: not published | $2.74 → ~$0.52 (-81%) only if implicit hits land | Short-prompt agents, bursty traffic, and anyone who needs a documented discount. The loser. |
| DeepSeek V4 Pro (peak) | $0.044 | none | not published | "a few hours to a few days" | $1.76 → $0.19 (-89%) | Regulated data, or anyone needing a guaranteed hit |
| Claude Sonnet 5.5, timestamp in system prompt | — | — | — | — | $2.72 → $2.56 (-6%) | Everyone. This is the bill engineers complain about. |
Workload, held identical across providers: a 20,000-token stable prefix (12,000 tokens of tool definitions, 8,000 of system prompt); 30 turns per session; each turn appends 1,500 new tokens (the last answer plus a tool result) and generates 400 output tokens; turns arrive inside each provider's cache lifetime. That is 1.30M input and 12,000 output tokens per session. Gemini's figure assumes implicit hits bill at its listed $0.20 context-caching rate, which its caching page does not confirm.
Why Doesn't Prompt Caching Save the Advertised 90%?
Because 90% is the discount on cached input tokens, not on your bill, and a real agent never caches all of its input. Prompt caching is the provider storing the processed form of a prompt prefix it has already seen, so a later request beginning with the identical bytes pays a reduced "cache read" price for that prefix instead of the full input rate.
Anthropic's launch post promised to "reduce costs by up to 90% and latency by up to 85% for long prompts." The same post's own table tells the fuller story: 90% for chatting with a cached 100K-token book, 86% for many-shot prompting, and 53% for a multi-turn conversation. The 90% is a best case on the workload that suits caching best. Agents are the multi-turn case.
Three things sit between the discount and your invoice:
- Output is never discounted. Every provider's cache price applies to input only. In our agent loop, output is under 5% of the uncached bill, so it barely matters. In a workload that writes long answers from short prompts it dominates. A 2,000-token prompt producing a 1,500-token answer on Claude Sonnet 5.5 spends $0.004 on input and $0.015 on output at Anthropic's $2/$10 rates. Caching the whole prompt perfectly saves under 20% of that call.
- The newest tokens are never cached on the turn they arrive. Each turn's fresh tool result is new to the provider. On Anthropic and on OpenAI's GPT-5.6-and-later models, that new material is written to cache at 1.25x the input rate so the next turn can read it.
- Writes cost more than reads save, per token. That is the trap. A miss on a cached prefix does not cost the plain input price; on the two biggest Western APIs it costs a quarter more.
In our loop, that arithmetic leaves Claude Sonnet 5.5 at 80% and GPT-6.1 Sol at 85% even with a perfect hit rate. Independent measurement lands in the same place. A PwC team that ran over 500 agent sessions on DeepResearch Bench, each with a 10,000-token system prompt, measured cost reductions of 41-80% across providers: 77.8% on Claude Sonnet 4.5, 79.3% on GPT-5.2, and just 38.3% on Gemini 2.5 Pro, with full-context caching (per-strategy results).
What Does Each Provider Charge to Write vs Read a Cache?
Anthropic and OpenAI charge a premium to write and a steep discount to read; Google and DeepSeek charge nothing extra to write, and that difference decides how much a cache miss hurts.
Anthropic (Claude). A 5-minute cache write costs 1.25x base input and a 1-hour write costs 2x; a read costs 0.1x, except 0.05x on Opus 5.5 and 0.025x on Fable 5.1. For Sonnet 5.5 that is $2.50 and $4.00 to write and $0.20 to read, per million tokens. Anthropic states the break-even plainly on the same page: caching "pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration." The multipliers stack with the Batch discount and with the 1.1x charge for US-only inference.
OpenAI. On GPT-5.6 and later, "cache writes cost 1.25x the standard, uncached input-token rate," and reads cost 0.1x, or 0.05x on GPT-6.1 Sol. Per the pricing page, GPT-6.1 Sol is $2.00 input, $0.10 cached input and $2.50 cache writes per million tokens up to 272K tokens, with the long-context tier doubling all of it. Earlier models had no write charge. If you migrated from GPT-5.4 and your bill did not fall, the new write fee is a candidate.
Google (Gemini). Implicit caching is "enabled by default for all Gemini 2.5 and newer models," and Google says it will "automatically pass on cost savings if your request hits caches." That caching page does not state the discount. The pricing page lists a context-caching price for Gemini 3.1 Pro Preview of $0.20 per million tokens at or under 200K and $0.40 above, plus $4.50 per million tokens per hour of storage for explicit caches. Gemini 3.8 Flash lists $0.075 through December 31, 2026, doubling to $0.15 on January 1, 2027, with storage at $0.50 per million tokens per hour, rising to $1.00 on the same date.
DeepSeek. Caching is on by default for every user with no code change, and there is no separate write price. V4 Pro bills $0.044 per million on a cache hit and $1.32 on a miss at peak, half that off-peak. Peak means 01:00-04:00 and 06:00-10:00 UTC on weekdays. A hit costs 3.3% of a miss, which is why DeepSeek gets closest to the 90% figure on the same workload.
What Are the Minimum Prefix and TTL Rules?
Every provider has a size floor below which nothing caches and a lifetime after which a cached prefix is gone. Both fail silently.
| Minimum prefix | Lifetime | Does a hit extend it? | |
|---|---|---|---|
| Claude Sonnet 5.5, Opus 5.5, Fable 5.1 | 512 tokens | 5 min default; 1 h optional | Yes, at no charge |
| Claude Haiku 4.5 | 4,096 tokens | same | Yes |
| GPT-5.6 and later | 1,024 tokens | 30 min, the only option | Window runs from last write or reuse |
| Gemini 3.1 Pro / 3.x Flash | 4,096 tokens | implicit: not published | not published |
| DeepSeek | not published | "hours to days", best-effort | not published |
Sources: Anthropic caching docs, OpenAI caching guide, Gemini caching docs, DeepSeek caching guide.
Four rules from the docs cause most of the "caching is on but nothing happened" tickets:
- Below the minimum, the request is processed uncached and nothing errors. Anthropic: "Prompts shorter than the minimum cannot be cached, even if marked with
cache_control." Moving an agent from Sonnet 5.5 (512) to Haiku 4.5 (4,096) can quietly turn caching off. - Anthropic's 5-minute clock starts when the request starts. If a response streams for four minutes, the next request has about one minute to arrive. Agents that wait on a slow tool or a human approval blow through that every time.
- Anthropic only looks back 20 blocks from a breakpoint. A long run of messages after your last
cache_controlmarker can find no hit at all. - OpenAI caches route by prefix to a machine. "Traffic above 15 requests per minute can lead to overflow routing" to machines that do not hold your cache. A busy shared prefix can see its hit rate fall as traffic rises.
The TTL rule has a measured price. Re-run our loop on Sonnet 5.5 with every turn arriving after the 5-minute window, and every turn rewrites the whole context at 1.25x: $3.36 a session, 24% more than not caching at all. Switch to the 1-hour write and the same slow loop costs $0.63, still a 77% saving. A July 2026 paper on cache keepalives for agentic workloads found Anthropic caches stone cold after 600 seconds idle (0 of 48 warm across three runs). Pinging roughly every 240 seconds kept them warm and cut the post-pause request's cost by up to 12.5x.
How Does a Moving System Prompt Destroy the Hit Rate?
A cache hit requires a byte-identical prefix, so any value that changes per request, placed early, invalidates everything after it, every turn. On Anthropic the hierarchy is tools, then system, then messages: changing a tool definition invalidates all three levels, and changing the system prompt invalidates the system and message caches.
The Manus team, whose agents run at an input-to-output ratio of about 100:1, calls KV-cache hit rate "the single most important metric for a production-stage AI agent." Their first rule is a stable prompt prefix, and their example is a timestamp at the start of the system prompt, because "even a single-token difference can invalidate the cache from that token onward." The PwC study gives the same advice: avoid "timestamps, datetime strings, session identifiers, or user-specific information" in the system prompt.
Here is what that costs on our workload, Claude Sonnet 5.5, per 30-turn session:
| Where the volatile value sits | Cache reads | Cache writes | Session cost | Saving vs no cache |
|---|---|---|---|---|
| Nowhere in the prefix (fixed) | 1.23M | 0.07M | $0.53 | 80% |
| First line of the system prompt | 0.35M (tools only) | 0.95M | $2.56 | 6% |
| Inside a tool description | 0 | 1.30M | $3.36 | -24% |
The pattern shows up in practitioner reports:
- A developer logged 3,412 calls over nine days at a 0% hit rate, caused by
datetime.now().isoformat()at the top of the system prompt. Fixing it reached 78%. Sorting the tools list so it serialised the same way every time reached 93%. - A September 24, 2026 openclaw issue traced cache-read drops rising from 7 to 54 a day at flat session volume to tool-set membership changes and a per-turn system-prompt suffix. The reporter estimated only about 25% of invalidations came from provider-side expiry and 75% were self-inflicted.
- An August 2026 Claude Code issue describes the case you cannot fix in your own code. In a ~950K-token session written to the 1-hour cache, reads collapsed to the 38,701-token system prefix every 33-61 seconds. Sixteen of 339 requests produced 51.8% of the window's cost. It had no maintainer response when we checked.
The silent invalidators to grep for: timestamps and dates, request or session IDs, the user's name or plan tier, tool lists built from a set or a dict without sorting, JSON serialised without stable key order, retrieved documents placed above the instructions, and switching thinking or effort settings mid-conversation.
What Does Caching Save on a Real Agent Loop?
On a well-built agent loop, 75-85% of the token bill is the realistic ceiling on Claude and OpenAI, about 89% on DeepSeek, and roughly half that on Gemini, going by the one independent benchmark. Our modelled loop gives the first figures. The PwC benchmark, which tested Gemini 2.5 Pro and does not say whether it used implicit or explicit caching, gives the last.
Where each provider lands, and who should avoid it:
Claude Sonnet 5.5: the most controllable. Explicit cache_control breakpoints (up to four), a 512-token floor, usage fields that separate reads from writes, and "cache hits are not deducted against rate limits." That last point is worth more than the price at volume, because it raises your effective throughput ceiling. The weakness is the write premium combined with a 5-minute default: a slow agent pays more than an uncached one. Don't pick it if your turns regularly wait on humans and you won't move the prefix to the 1-hour write.
GPT-6.1 Sol: the cheapest frontier cache read. At $0.10, its read is half Sonnet 5.5's at the same $2 input rate, and its 30-minute window forgives slow tools. It wins our loop among the Western APIs, $0.41 against $0.53. But caching is something OpenAI does for you, not something you pin. prompt_cache_key helps routing and cache accounting, while overflow above ~15 requests a minute per prefix is documented behaviour. Don't pick it if a finance owner needs a guaranteed hit rate in a contract or a forecast.
DeepSeek V4 Pro: closest to the brochure. With no write fee and a hit at 3.3% of a miss, a broken cache costs you the saving but never a penalty. That makes it the most forgiving model in this list. The docs call the cache "best-effort" and say it "does not guarantee a 100% cache hit rate." The peak/off-peak split also means the same job can cost twice as much depending on the hour. Don't pick it for data your security team would not send to a provider in mainland China. We covered that hedge in DeepSeek Will Raise Prices.
Gemini 3.1 Pro: the loser for agent caching. The 4,096-token floor shuts out the short-prefix agents that make up most internal tooling. The implicit discount is unquantified on the caching page itself. The explicit alternative bills storage by the hour whether or not anyone calls it. And Gemini 2.5 Pro saved only 38.3% in the one independent agent benchmark, against 77.8-79.3% for Claude Sonnet 4.5 and GPT-5.2. Gemini 3.1 Pro is a newer model and may do better, but nobody independent has published that, and it is still labelled Preview. Don't pick it on caching economics. Pick it for capability, then measure your own hit rate.
When Is Restructuring Your Prompts Not Worth It?
When input tokens are a small share of the bill, when the stable prefix is under the provider's floor, or when each prefix is read fewer times than it takes to repay the write. Check these before you spend a sprint on cache engineering:
- Output-heavy workloads. Your maximum saving is roughly the input share of the bill times the read discount. If input is 20% of spend, the most caching can return is about 18%.
- Prefixes below the minimum. A 3,000-token prompt on Gemini or Haiku 4.5 cannot cache at all. The fix there is a model with a lower floor, not prompt surgery.
- Low reuse inside the TTL. On Anthropic's 5-minute tier you need at least one read per write; on the 1-hour tier, two. A batch job that touches each document once gets nothing, and the Batch API's 50% discount is the better lever.
- Per-user content that must come first. If permissions or personalisation force user data above the instructions, every user has a private prefix and the shared cache is gone. We hit the same wall in Long Context vs RAG Cost.
- You already compress the prompt. A July 2026 study found query-aware compression made a τ-bench retail agent 40.1% more expensive than doing nothing, because it rewrote the cached prefix on every call. Compress the volatile tail, never the stable head.
Which Criteria Actually Predict Regret?
The decision rarely turns on the read price. It turns on four numbers you can pull from logs this week.
- Hit rate, per workload. Compute it as cache-read tokens over total input tokens:
cache_read_input_tokenson Anthropic,input_tokens_details.cached_tokenson OpenAI,prompt_cache_hit_tokenson DeepSeek. Under 70% on an agent loop means a volatile prefix, not a pricing problem. - Write-to-read ratio. On Anthropic and OpenAI, cache-write tokens above about 10% of cache reads means you are re-writing what you should be reading.
- Median gap between turns. If it is over five minutes, Claude's default tier is the wrong tier.
- Input share of spend. Under 30% and caching is a rounding error, whatever the provider.
What changes the answer: a cut to a model's read multiplier (Anthropic already prices reads below 0.1x on Opus 5.5 and Fable 5.1), a vendor adding or removing a write fee (OpenAI added one with GPT-5.6), or a gateway that rewrites prompts. Gateways that inject headers or reorder tools can break a cache your application built correctly. Check any gateway you route through before blaming the provider.
What to Do About It
This Week:
- Pull 30 days of usage per workload and compute hit rate and write-to-read ratio from the fields above. An observability tool or LiteLLM's spend logs will usually have them. If you can't produce the numbers, that is the first finding.
- Grep every prompt template for
now(),uuid,request_id, user names and dates. Move each one to the last user message. - Sort tool definitions and serialise JSON with stable key order. It is a one-line fix that a developer measured moving a hit rate from 78% to 93%.
This Month:
- Alert on hit rate, not just spend. A drop from 90% to 10% more than quadruples the Claude cost of our loop, and it happens on a deploy, not over a quarter.
- Measure your median inter-turn gap. Where it exceeds five minutes on Claude, move the stable prefix to the 1-hour write and let it pay for itself after two reads.
- A/B the same agent on two providers with identical prefixes, and assert that cache-read tokens are non-zero before trusting either result.
Before Renewal:
- Price any usage commitment at your measured hit rate, not the vendor's. A commitment sized on uncached tokens is oversized once you fix the prefix, and that gap is leverage. Our per-provider cost teardown walks through pricing the task rather than the rate.
The Bottom Line
Prompt caching is a real discount, offered on every major API without a contract. It is also the one most often lost to a line of your own code. The 90% headline describes a cached input token. Your invoice describes a workload, and on a well-built agent that workload saves 80-89%, not 90%. On a badly built one it saves single digits, or costs more.
This repeats an older pattern. CDNs promised the same economics two decades ago, and teams that put a session ID in every URL paid full origin cost while assuming the edge had their back. The fix was never a better CDN. It was a stable key.
Find the timestamp before you shop for a cheaper model.
Continue Reading
- Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.
- One PR Billed 156M Tokens. Cap the Reads, Not the Rate.
- Your AI Router Is Trading a 10x Discount for a 2.5x One
- Long Context vs RAG Cost: Cached, Each Answer Still Costs 5x
- Inference Cost per Million Tokens: Price the Task, Not the Rate
- Best LLM Gateways for Cost Control: Self-Host First
