Topic

model routing

Every THE D[AI]LY BRIEF article on model routing — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

Claude Haiku 5.5

Claude Haiku 5.5 Charges 5x Once a Prompt Passes 100K Tokens

Claude Haiku 5.5 matches GPT-6 Luna at $0.10/$0.50 per million tokens and beats it on Anthropic's evals, but prompts over 100K tokens cost 5x and its tokenizer counts about 30% more tokens than Haiku 4.5. Short-prompt classification and routing get the big saving; compaction and long-context work get much less.

October 7, 2026 · 11 min read
Clef

Cloudflare's Clef Beats Jev at Routing and Loses at Judgment

Cloudflare's Apache 2.0 Clef beats TypeSafe's Jev on intent and tool-call benchmarks, trails it on When2Call, and costs about 5.7x more per input token hosted. Amazon's Strands Decider 2B runs on a consumer GPU but scores 0.505 on hard tasks.

October 3, 2026 · 11 min read
Claude Sonnet 5.5

Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score

Claude Sonnet 5.5 lists at half Opus 5.5's token price, but Artificial Analysis measured it at $7.60 per task at max effort against $3.46 for Opus at xhigh for the same score. It saves money only at low and high effort, and five breaking changes sit behind the model-ID swap.

September 28, 2026 · 10 min read
GitHub Copilot

GitHub's Router Bills Every Leg. Find Your Break-Even.

GitHub's HydraFusion router bills every model it invokes at standard rates, and two of its three patterns fire two or three models per turn. The measured saving swings from 36% to 67% across just three benchmarks — and the one that looks most like enterprise work saved least.

September 5, 2026 · 12 min read
prompt caching

Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.

Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.

September 1, 2026 · 15 min read
AI coding agents

Haiku Burned More Tokens Than Sonnet. Spec It in Code.

A controlled 90-trial experiment found Claude Haiku 4.5 spent 735K tokens where Sonnet 4.6 spent 640K, for a result 1.9 points worse. Downgrading a coding agent to a cheap tier saves less than the rate card implies, varies fivefold by vendor, and only holds up if you replace prose design docs with machine-checkable contracts.

August 25, 2026 · 12 min read
AI spending

100% of CIOs Budgeting for AI. Half Already Blew Their Budgets.

RBC's CIO survey shows 100% budgeting for AI, 90% increasing spend, and 91% creating entirely new budgets. But underneath: Uber burned its annual AI budget in 4 months, Microsoft is canceling Claude Code licenses, and 73% of enterprises exceeded projections. The gap between budget intent and budget reality is the AI FinOps crisis of 2026. Spend health assessment and model routing matrix inside.

June 27, 2026 · 17 min read