Topic

cache hit rate

Every THE D[AI]LY BRIEF article on cache hit rate — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

prompt caching

Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.

Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.

September 1, 2026 · 15 min read
prompt caching

Your AI Router Is Trading a 10x Discount for a 2.5x One

Manifest killed its four-tier LLM router after four months and 7,000 users, and the arithmetic explains why: cache reads bill at 10% of base input, so routing an agent step to a model 2.5x cheaper makes it 3.5x more expensive. Route at the session boundary, not the request.

August 1, 2026 · 15 min read