Long Context vs RAG Cost: Cached, Each Answer Still Costs 5x

Prompt caching cuts a 400K-token long-context answer to $0.085 on Claude Sonnet 5.5, but RAG still answers for $0.017. The crossover formula, the accuracy-at-depth evidence and the latency users notice.

By Rajesh Beri·October 1, 2026·14 min read
Share:
A tall stack of thick printed binders on a conference table beside a single slim index-card box, with a desk calculator and a printed invoice between them, lit by late-afternoon office light.

Illustration generated using AI

Prompt caching made long context cheap enough to tempt you into deleting your retrieval layer — and not cheap enough to justify it at scale. On a 400,000-token corpus, a cached long-context answer on Claude Sonnet 5.5 costs about $0.085; the same answer through a retrieval pipeline costs about $0.017. That 5x gap is noise at 100 queries a day and roughly $200,000 a month at 100,000. The verdict: below ~1,000 queries a day, on a corpus that fits, changes rarely and has one permission boundary, skip RAG and use long context with caching. Above a few thousand a day, or the moment users have different access rights, build retrieval — and put it on pgvector.

All prices below were checked against each vendor's live pricing page on October 1, 2026.

Architecture (400K-token corpus) Cost per answer 100 queries/day 1,000/day 100,000/day Who should NOT pick it
Long context, uncached, Claude Sonnet 5.5 $0.805 $80.50 $805 $80,500 Anyone. This is the loser.
Long context, cached, Claude Sonnet 5.5 (1h TTL) $0.085 $10.05 $86.91 $8,541 Teams whose corpus changes more than a few times an hour
Long context, cached, GPT-6.1 Sol (>272K tier) $0.088 $10.75 $90.22 $8,832 Anyone who can't tolerate an unguaranteed cache hit
Long context, explicit cache, Gemini 3.1 Pro (>200K tier) $0.170 + $1.80/hour storage $36.42 $189.24 $16,999 Bursty, low-volume workloads — storage bills while idle
RAG on pgvector + Sonnet 5.5 (6K-token prompt) $0.017 $1.70 $17 $1,700 Teams with no one to own an index and an eval set

Workload: a 400,000-token corpus, a 200-token question, a 500-token answer, traffic across a 10-hour business day, one cache write per day. RAG retrieves ~5,000 tokens of chunks plus a 1,000-token system prompt. Token prices only; RAG's fixed cost is handled in the crossover formula below.


What Does a Long-Context Answer Actually Cost?

An uncached long-context answer costs the full corpus at the input rate, every single time. On Claude Sonnet 5.5 that is $2 per million input tokens and $10 per million output, per Anthropic's pricing page, so 400,200 input tokens is $0.80 before the model writes a word. Anthropic bills its full 1M-token window at standard rates — "a 900k-token request is billed at the same per-token rate as a 9k-token request," per the same page.

The other two vendors charge a long-context surcharge, and a 400K corpus trips it. OpenAI's pricing doubles GPT-6.1 Sol from $2 to $4 per million input tokens (and $10 to $15 output) once a request exceeds 272K input tokens. Google's Gemini pricing doubles Gemini 3.1 Pro Preview from $2 to $4 input and lifts output from $12 to $18 above 200K tokens. A corpus that sits just over those lines costs twice what one just under them does — trimming 130K tokens of boilerplate out of a 400K corpus can halve your GPT bill before you touch architecture.

One caveat that moves every number: Anthropic says Claude 4.7 and later models use a tokenizer that "produces approximately 30% more tokens for the same text" (pricing page). The table uses 400K tokens as each vendor counts them. If your corpus is 400K tokens on OpenAI's tokenizer, expect it to count higher on Claude.

How Much Does Prompt Caching Change the Long-Context Number?

Prompt caching cuts the long-context answer by roughly 90% — from $0.805 to $0.085 on Sonnet 5.5 — and that is the whole reason this debate is back. Prompt caching is the provider storing the processed form of an unchanged prompt prefix, so later requests that start with the same prefix pay a reduced "cache read" rate instead of full input price.

The three implementations differ in ways that matter to the bill:

  • Claude Sonnet 5.5 charges $0.20 per million tokens for a cache hit, $2.50 for a 5-minute write and $4 for a 1-hour write (pricing). Hits refresh the TTL, and Anthropic's caching docs state "cache hits are not deducted against your rate limit" — which matters more than the price once you push 400K tokens per request. At 100 queries spread over a 10-hour day, requests arrive about every six minutes, so the 5-minute cache would keep expiring; pay the $1.60 one-hour write instead.
  • GPT-6.1 Sol caches automatically — "prompt caching is enabled by default," per OpenAI's caching guide — and on GPT-5.6 and later charges cache writes at 1.25x the uncached input rate, with a 30-minute window after the last write or reuse. Above 272K tokens a cached token costs $0.20 per million (pricing). The catch is that you are not buying a guaranteed hit; you are hoping for one.
  • Gemini 3.1 Pro offers implicit caching, "enabled by default for all Gemini 2.5 and newer models" with a 4,096-token minimum (caching docs), plus explicit caches priced at $0.40 per million tokens read above 200K and $4.50 per million tokens per hour of storage (pricing). A 400K explicit cache costs $1.80 an hour whether anyone asks a question or not.

Gemini 3.1 Pro is the loser on this workload, and not narrowly: double the cache-read rate of the other two above its 200K threshold, plus a storage meter that runs while your users sleep. Under 200K tokens its cache read drops to $0.20 and the comparison tightens — but it also remains a Preview model, which is its own reason not to anchor a production cost model to it.

What Does the Same Answer Cost Through RAG?

A RAG answer costs about $0.017 in tokens, and the retrieval infrastructure rounds to zero at this corpus size — the real cost is the people. RAG (retrieval-augmented generation) is the pattern of embedding the corpus into a vector index, retrieving only the chunks relevant to each question, and sending those to the model instead of everything.

The token side: ~6,000 input tokens at Sonnet 5.5's $2 rate is $0.012, plus the same $0.005 of output. Embedding the whole 400K corpus with OpenAI's text-embedding-3-small costs $0.008 at $0.02 per million tokens; re-embedding it every night costs less than the coffee you drink while reading the invoice.

The index side: pgvector is an open-source Postgres extension, so if you already run Postgres the incremental cost of a few thousand chunks is effectively nothing. A managed alternative like Pinecone sets a $50/month minimum on its Standard plan, with read units at $16-$18 per million and storage at $0.33/GB-month (Pinecone pricing). Even if every query consumed a full read unit, 3 million queries a month would add roughly $50 — real money only in the sense that it is a line item. We made the case for staying on Postgres in Pinecone vs Weaviate vs pgvector, and nothing here changes it.

What RAG really costs is an owner: someone who tunes chunking, maintains an eval set, and fixes the retrieval miss that a customer escalates. That is the fixed cost F in the formula below, and it is why RAG loses at low volume. Our 10M-tokens-a-day pipeline teardown puts the infrastructure around it in dollars.

Who should not pick RAG: a team with no one who will own retrieval quality after launch. An unowned index degrades silently, and silent wrong answers are worse than expensive right ones.


Where Is the Crossover Point?

RAG pays for itself when the per-answer token savings, multiplied by monthly volume, exceed what it costs you to own a retrieval system. As a formula:

N* = F ÷ (C × r_cache − R × r_in)

N* = monthly queries at which RAG becomes cheaper · F = fixed monthly cost of owning retrieval (engineering time, infra, eval upkeep) · C = corpus tokens · r_cache = cache-read price per token · R = RAG prompt tokens · r_in = standard input price per token

Output cost cancels because both architectures write the same answer. Add W × r_write × (corpus changes per month) to the long-context side if your corpus changes — every edit to the prefix invalidates the cache and forces a fresh write, $1.60 a time on Sonnet 5.5 at 400K.

Plugging in Sonnet 5.5 at 400K: C × r_cache = $0.08, R × r_in = $0.012, so each answer saves $0.068 under RAG.

  • If owning retrieval costs you F = $5,000 a month (a fraction of an engineer plus infra), N* ≈ 73,500 queries a month, or ~2,450 a day.
  • If it costs F = $20,000 (a dedicated engineer), N* ≈ 294,000 a month, or ~9,800 a day.
  • Without caching, the saving per answer is $0.788, and the $5,000 crossover collapses to ~211 queries a day.

That is the number to take to your architecture review: caching moves the crossover about 11x. Every "RAG is dead" post you read is using the cached number; every "long context is unaffordable" post is using the uncached one. Both are right about the wrong workload. Our earlier fine-tuning vs RAG break-even uses the same F-over-delta logic for a different decision.

How Accurate Is a Model at Depth in a Large Context?

Accuracy degrades well before the advertised window is full, and 400K tokens is deep enough to feel it. Anthropic's own context-window documentation says it plainly: "As token count grows, accuracy and recall degrade, a phenomenon known as context rot."

The independent evidence agrees:

  • NoLiMa (ICML 2025), which removes literal keyword overlap between question and answer, found 11 of 13 models fell below 50% of their short-context score at just 32K tokens; GPT-4o dropped from 99.3% to 69.7%.
  • Chroma's Context Rot study (July 2025) tested 18 models and found performance "varies significantly as input length changes, even on simple tasks"; on LongMemEval, focused ~300-token prompts beat the full ~113K-token history.
  • Lost in the Middle (TACL) documented the U-shaped curve: facts at the start or end of the context are found; facts in the middle are missed.
  • On OpenAI's MRCR v2 8-needle test, a March 2026 compilation of vendor and third-party scores shows Claude Opus 4.6 at 93.0% at 256K falling to 76.0% at 1M, and Gemini 3 Pro at 45.4% at 256K falling to 24.5% at 1M. Those are prior-generation models; re-run the test on the model you are buying, because the shape of the curve is the part that has held.

Steel-man for long context: retrieval misses too, and it misses quietly. A June 2026 study of document-grounded answers (Hamilton et al.) found long-context prompting 73.1% correct versus 65.4% for semantic RAG — at 26 times the per-query token cost. A practitioner test of Kimi K3 on a 127K-token corpus (Towards Data Science, Aug 2026) found long context gave complete answers on all 12 questions while RAG scored 0.83/2.0 on completeness — mostly on corpus-wide questions that no top-5 chunk set can answer.

So the honest accuracy rule: long context wins on questions that span the whole corpus ("summarise every change across these contracts"); RAG wins on needle questions in a corpus deep enough to rot. If your traffic is mostly the second kind, the cheaper architecture is also the more accurate one.

Will Users Notice the Latency Difference?

Yes — uncached long context is slow enough to change the product, and caching closes most but not all of the gap. In the Kimi K3 test above, the long-context path averaged 111 seconds per answer versus 11 for RAG, and 273.7 seconds versus 46.3 on the hardest corpus-wide questions (Towards Data Science). That model and harness are not yours, but the 10x ratio is the order of magnitude to plan for.

Caching attacks the prefill, which is most of time-to-first-token. Anthropic's published example — chatting with a 100,000-token book — drops time-to-first-token from 11.5s to 2.4s, a 79% cut (Claude blog). At 400K, scale your expectations up, not down. And a cache miss reverts to the full prefill: the first user after an idle hour, and every user after a corpus edit, gets the slow path.

A RAG prompt of 6,000 tokens has almost no prefill to speak of. If your product promise is a chat answer that starts streaming in under two seconds, long context at 400K cannot keep it reliably, cached or not.


Which Should You Pick? The Criteria That Predict Regret

Pick long context with caching when all four of these hold; build RAG when any one of them breaks.

  1. Volume stays under your N*. Compute it with your own F. At typical numbers that is ~1,000-2,500 queries a day on Sonnet 5.5 or GPT-6.1 Sol.
  2. The corpus changes rarely. Hourly edits mean hourly $1.60 cache writes and a slow first answer after each one. A knowledge base that updates nightly is fine; a ticket queue is not.
  3. Everyone may see everything. This is the criterion architects forget. Long context with a shared cache only works if every user is entitled to the whole corpus. Per-user permissions mean per-user prefixes, which means no shared cache, which means you are back on the $0.805 row. Our RAG build-vs-buy analysis covers why permission sync is the real work.
  4. The corpus fits comfortably under the surcharge line. Under 200K tokens on Gemini, under 272K on GPT. Claude has no line, but context rot does.

What changes the answer: a cache-read price cut (Anthropic already prices Opus 5.5 hits at 0.05x base and Fable 5.1 at 0.025x, per its pricing page — if that reaches Sonnet, Δ shrinks and N* rises); a corpus that grows past 1M tokens, which ends the debate; or a compliance requirement to cite the exact source passage, which RAG's chunk provenance gives you for free.

The recommendation for each option:

  • Claude Sonnet 5.5 cached — the default long-context choice: flat pricing across 1M, explicit cache control, cache hits outside your rate limit. Don't pick it if your corpus mutates more than a few times an hour.
  • GPT-6.1 Sol cached — nearly the same per-answer cost with zero cache engineering. Don't pick it if you need to prove a hit rate to finance; you can measure it, but you can't pin it.
  • Gemini 3.1 Pro explicit cache — the loser above 200K tokens. Don't pick it for a 400K corpus with business-hours traffic; reconsider it under 200K with steady 24-hour load where the storage fee amortises.
  • RAG on pgvector — the answer above a few thousand queries a day and the only answer with per-user permissions. Don't pick it without a named owner and an eval set.

This Week: count your corpus in each vendor's tokenizer and your real daily query volume; most teams guess both wrong by 2x. This Month: put F on paper — the engineering hours retrieval actually consumes — and compute N*. If you are below it, prototype long context with a 1-hour cache and log cache_read_input_tokens on every call. Before Q1 Close: run 50 needle questions and 20 corpus-wide questions against both architectures on your own documents. Price the wrong answers, not just the right ones.

The Bottom Line

Long context did not kill RAG; caching repriced the threshold where RAG earns its keep, from a couple of hundred queries a day to a couple of thousand. That is a real shift — most internal assistants live below it, and for them the retrieval layer is now an engineering cost with no business case. But the enterprise workloads that matter most have high volume, mutable corpora and permissions, and every one of those drags you back toward retrieval.

Treat this like the buy-versus-colocate decision of the last infrastructure cycle: the convenient option wins small, the engineered one wins at scale, and the mistake is building for the scale you don't have. Price the answer, not the window.

Continue Reading

Share:

Frequently Asked Questions

Is long context cheaper than RAG with prompt caching?

Per answer, no. On a 400K-token corpus, a cached long-context answer on Claude Sonnet 5.5 costs about $0.085 versus about $0.017 for RAG (October 2026 list prices). Long context is cheaper overall only when volume is low enough that RAG's fixed engineering cost outweighs the per-answer saving.

At what query volume does RAG become cheaper than long context?

Use N* = F / (C x cache-read price - R x input price), where F is the monthly cost of owning retrieval, C the corpus tokens and R the RAG prompt tokens. With Sonnet 5.5, a 400K corpus and F = $5,000/month, the crossover is about 2,450 queries a day; without caching it falls to about 211.

Does accuracy drop in a very large context window?

Yes. Anthropic's own docs call it context rot. NoLiMa found 11 of 13 models fell below half their short-context score at 32K tokens, and MRCR v2 scores fall sharply between 256K and 1M. Long context still wins on questions that span the whole corpus.

Why do per-user permissions push you toward RAG?

A shared prompt cache only works if every user may see the whole corpus. Different access rights mean different prompt prefixes per user, which defeats the cache and puts you back on uncached long-context pricing, about $0.805 per answer at 400K tokens on Sonnet 5.5.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →