Topic

model selection

Every THE D[AI]LY BRIEF article on model selection — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

Real-SWE

Fable Cost $18 a Fix on Private Code. Flash Cost $8.

On Specific Labs' Real-SWE benchmark of licensed private codebases, Fable 5.1 resolved the most tasks but cost about $17.94 per fix to Gemini 3.8 Flash's $8.01 — and each model fails in its own way, so your review gate should follow the agent you pick.

September 13, 2026 · 14 min read
MCP

150 MCP Servers Never Started. No Benchmark Shows That.

An unrepaired probability sample of 400 MCP servers found only 48.8% complete an initialize handshake, and 37.5% never start at all. Meanwhile 68.8% of BFCL v4's extracted tool definitions and 85.6% of UltraTool's are exact repeats, against 0.4% for real MCP tools.

September 12, 2026 · 11 min read
agent evaluation

Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.

The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.

September 10, 2026 · 13 min read
prompt caching

Your AI Router Is Trading a 10x Discount for a 2.5x One

Manifest killed its four-tier LLM router after four months and 7,000 users, and the arithmetic explains why: cache reads bill at 10% of base input, so routing an agent step to a model 2.5x cheaper makes it 3.5x more expensive. Route at the session boundary, not the request.

August 1, 2026 · 15 min read
GPT-5.6

GPT-5.6: 3 Models, 30x Price Spread, 1 Enterprise Decision

OpenAI just split GPT-5.6 into three models — Sol ($5/$30), Terra ($2.50/$15), and Luna ($1/$6) per million tokens. Sol beats Anthropic's restricted Mythos on TerminalBench. Terra matches GPT-5.5 at half the cost. The release is limited to ~20 organizations under a new U.S. government review process, but the pricing and benchmarks are public. Here's how to classify your workloads across all three tiers and prepare for migration before general availability.

June 26, 2026 · 17 min read