Claude Haiku 5.5Claude Haiku 5.5 Charges 5x Once a Prompt Passes 100K Tokens
Claude Haiku 5.5 matches GPT-6 Luna at $0.10/$0.50 per million tokens and beats it on Anthropic's evals, but prompts over 100K tokens cost 5x and its tokenizer counts about 30% more tokens than Haiku 4.5. Short-prompt classification and routing get the big saving; compaction and long-context work get much less.
October 7, 2026 · 11 min readIBM BobIBM's Air-Gapped Bob Drops Claude for Nemotron and Laguna
Self-hosted IBM Bob swaps SaaS Bob's Claude, Mistral and Granite routing for one local model, Nemotron or Laguna. IBM has published no evals, GPU sizing or price for it.
October 4, 2026 · 9 min readClefCloudflare's Clef Beats Jev at Routing and Loses at Judgment
Cloudflare's Apache 2.0 Clef beats TypeSafe's Jev on intent and tool-call benchmarks, trails it on When2Call, and costs about 5.7x more per input token hosted. Amazon's Strands Decider 2B runs on a consumer GPU but scores 0.505 on hard tasks.
October 3, 2026 · 11 min readCloudflare AI GatewayCloudflare's Auto Router Fails 4x as Often as Opus on Its Own Test
Cloudflare's Auto Router solved 86.6% of its own benchmark tasks against Opus 5.5's 96.6%, at 40% of the cost per success. The saving only holds where a failed task costs under about 13 cents.
October 1, 2026 · 11 min readGPT-6.1 SolGPT-6.1 Sol's $5.47 Task Scores 11 Points Below Astra's $23.80
GPT-6.1 Sol costs a fifth of GPT-6 Astra per token and ties it on coding, but trails it by 11 points on Terminal-Bench Science. It still wins on cost per solved task — if your failures are cheap.
September 29, 2026 · 8 min readClaude Sonnet 5.5Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score
Claude Sonnet 5.5 lists at half Opus 5.5's token price, but Artificial Analysis measured it at $7.60 per task at max effort against $3.46 for Opus at xhigh for the same score. It saves money only at low and high effort, and five breaking changes sit behind the model-ID swap.
September 28, 2026 · 10 min readTypeSafe JevTypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
TypeSafe's Jev answers classification calls for 12-27x less than Claude Haiku 4.5, but the first independent tests show its accuracy depends on how you split the question and its probabilities need recalibrating per question.
September 20, 2026 · 13 min readmodel routingModel Router Buyer's Guide: Buy Failover, Not Judgment
A model router sells two things: mechanical failover, which works, and a learned classifier that picks your model, which the neutral benchmarks say does not. Buy the first, build the second, and plan on 35% savings rather than 60%.
September 11, 2026 · 20 min readGitHub CopilotGitHub's Router Bills Every Leg. Find Your Break-Even.
GitHub's HydraFusion router bills every model it invokes at standard rates, and two of its three patterns fire two or three models per turn. The measured saving swings from 36% to 67% across just three benchmarks — and the one that looks most like enterprise work saved least.
September 5, 2026 · 12 min readprompt cachingAnthropic Cut Cache Reads 75%. Opus 5 Still Undercuts It.
Anthropic cut Claude Fable 5.1 cache reads 75% to $0.25/MTok, half of Opus 5's rate on a model with twice the base input price. But Opus 5 stays cheaper until cache reads pass a third of your bill, and Anthropic's own savings ceiling of 45% falls short of the 50% crossover.
September 1, 2026 · 15 min readAI coding agentsHaiku Burned More Tokens Than Sonnet. Spec It in Code.
A controlled 90-trial experiment found Claude Haiku 4.5 spent 735K tokens where Sonnet 4.6 spent 640K, for a result 1.9 points worse. Downgrading a coding agent to a cheap tier saves less than the rate card implies, varies fivefold by vendor, and only holds up if you replace prose design docs with machine-checkable contracts.
August 25, 2026 · 12 min readAI spending100% of CIOs Budgeting for AI. Half Already Blew Their Budgets.
RBC's CIO survey shows 100% budgeting for AI, 90% increasing spend, and 91% creating entirely new budgets. But underneath: Uber burned its annual AI budget in 4 months, Microsoft is canceling Claude Code licenses, and 73% of enterprises exceeded projections. The gap between budget intent and budget reality is the AI FinOps crisis of 2026. Spend health assessment and model routing matrix inside.
June 27, 2026 · 17 min read