model routingModel Router Buyer's Guide: Buy Failover, Not Judgment
A model router sells two things: mechanical failover, which works, and a learned classifier that picks your model, which the neutral benchmarks say does not. Buy the first, build the second, and plan on 35% savings rather than 60%.
September 11, 2026 · 20 min readGitHub CopilotGitHub's Router Bills Every Leg. Find Your Break-Even.
GitHub's HydraFusion router bills every model it invokes at standard rates, and two of its three patterns fire two or three models per turn. The measured saving swings from 36% to 67% across just three benchmarks — and the one that looks most like enterprise work saved least.
September 5, 2026 · 12 min readprompt cachingAnthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.
Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.
September 1, 2026 · 15 min readAI coding agentsHaiku Burned More Tokens Than Sonnet. Spec It in Code.
A controlled 90-trial experiment found Claude Haiku 4.5 spent 735K tokens where Sonnet 4.6 spent 640K, for a result 1.9 points worse. Downgrading a coding agent to a cheap tier saves less than the rate card implies, varies fivefold by vendor, and only holds up if you replace prose design docs with machine-checkable contracts.
August 25, 2026 · 12 min readAI spending100% of CIOs Budgeting for AI. Half Already Blew Their Budgets.
RBC's CIO survey shows 100% budgeting for AI, 90% increasing spend, and 91% creating entirely new budgets. But underneath: Uber burned its annual AI budget in 4 months, Microsoft is canceling Claude Code licenses, and 73% of enterprises exceeded projections. The gap between budget intent and budget reality is the AI FinOps crisis of 2026. Spend health assessment and model routing matrix inside.
June 27, 2026 · 17 min read