Topic

model routing

Every THE D[AI]LY BRIEF article on model routing — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

GitHub Copilot

GitHub's Router Bills Every Leg. Find Your Break-Even.

GitHub's HydraFusion router bills every model it invokes at standard rates, and two of its three patterns fire two or three models per turn. The measured saving swings from 36% to 67% across just three benchmarks — and the one that looks most like enterprise work saved least.

September 5, 2026 · 12 min read
prompt caching

Anthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.

Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.

September 1, 2026 · 15 min read
AI coding agents

Haiku Burned More Tokens Than Sonnet. Spec It in Code.

A controlled 90-trial experiment found Claude Haiku 4.5 spent 735K tokens where Sonnet 4.6 spent 640K, for a result 1.9 points worse. Downgrading a coding agent to a cheap tier saves less than the rate card implies, varies fivefold by vendor, and only holds up if you replace prose design docs with machine-checkable contracts.

August 25, 2026 · 12 min read
AI spending

100% of CIOs Budgeting for AI. Half Already Blew Their Budgets.

RBC's CIO survey shows 100% budgeting for AI, 90% increasing spend, and 91% creating entirely new budgets. But underneath: Uber burned its annual AI budget in 4 months, Microsoft is canceling Claude Code licenses, and 73% of enterprises exceeded projections. The gap between budget intent and budget reality is the AI FinOps crisis of 2026. Spend health assessment and model routing matrix inside.

June 27, 2026 · 17 min read
model routing Articles | THE D*AI*LY BRIEF