vendor benchmarksVendor AI Benchmarks: Only Numbers With Logs Belong in the Deck
A vendor-run benchmark is a screening signal. Business cases should cite independent runs with per-task logs, led by Epoch AI and Artificial Analysis, plus a 50-task replication on your own data.
October 7, 2026 · 16 min readGemini 4 ArgonGemini 4 Argon Leads Vals at $15.68 a Task With No GA Date
Gemini 4 Argon tops the Vals Index at $15.68 per test, cheaper than Claude Sonnet 5.5 and Opus 5.5 even at its $4/$20 standard price. But access is limited to Fairwind security teams and trusted testers, with no GA date.
September 30, 2026 · 7 min readSalesforce KoaSalesforce Koa vs Claude vs Gemini: Its Paper Only Beats GPT-4.1
Salesforce's own Koa paper shows it beating GPT-4.1 but trailing frontier models, and Koa has no GA date outside the U.S. and no price. Choose Agentforce models per sub-agent on your own held-out cases.
September 27, 2026 · 12 min readAstra for LawAstra for Law Scored 54% on a Set Vals Never Publishes
OpenAI scored Astra for Law at 54% on Vals AI's private validation set, against its own GPT-6 Astra. On Vals' public leaderboard, three general frontier models already sit at 55.29% and Astra for Law is not listed.
September 22, 2026 · 9 min readGrok 4.7Grok 4.7 Kept Its $2/$6 Price and Can Still Double Cost per Task
Grok 4.7's per-token price didn't move. Its per-task cost did: up 47% at matched effort on Artificial Analysis's index, down 10% on Cursor's own benchmark. Price it on your workload.
September 22, 2026 · 11 min readTypeSafe JevTypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways
TypeSafe's Jev answers classification calls for 12-27x less than Claude Haiku 4.5, but the first independent tests show its accuracy depends on how you split the question and its probabilities need recalibrating per question.
September 20, 2026 · 13 min readagent harnessYour Bake-Off Ranked Harnesses, Not Models. Re-Run It.
A 176-setting controlled ablation shows the coding-agent harness moves success and cost more than the gap between the models being compared — and that one tool-surface flag reverses which model wins. A model swap is not a drop-in.
September 20, 2026 · 12 min readReal-SWEFable Cost $18 a Fix on Private Code. Flash Cost $8.
On Specific Labs' Real-SWE benchmark of licensed private codebases, Fable 5.1 resolved the most tasks but cost about $17.94 per fix to Gemini 3.8 Flash's $8.01 — and each model fails in its own way, so your review gate should follow the agent you pick.
September 13, 2026 · 14 min readMCP150 MCP Servers Never Started. No Benchmark Shows That.
An unrepaired probability sample of 400 MCP servers found only 48.8% complete an initialize handshake, and 37.5% never start at all. Meanwhile 68.8% of BFCL v4's extracted tool definitions and 85.6% of UltraTool's are exact repeats, against 0.4% for real MCP tools.
September 12, 2026 · 11 min readmodel routingModel Router Buyer's Guide: Buy Failover, Not Judgment
A model router sells two things: mechanical failover, which works, and a learned classifier that picks your model, which the neutral benchmarks say does not. Buy the first, build the second, and plan on 35% savings rather than 60%.
September 11, 2026 · 20 min readagent evaluationOnly 3 of 36 Model Gaps Were Real. Size Your Eval Set.
The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.
September 10, 2026 · 13 min readagent evaluationDeepSeek Won Solo. Gemini Won the Arena. Run Both Rounds.
ERPBench ran 100 identical ERP problems through six model families twice — solo against fixed opponents, then all six competing in one market. DeepSeek won the first, Gemini the second, and the two agreed on 21 of 100 tasks.
September 7, 2026 · 12 min readprompt cachingYour AI Router Is Trading a 10x Discount for a 2.5x One
Manifest killed its four-tier LLM router after four months and 7,000 users, and the arithmetic explains why: cache reads bill at 10% of base input, so routing an agent step to a model 2.5x cheaper makes it 3.5x more expensive. Route at the session boundary, not the request.
August 1, 2026 · 15 min readGPT-5.6GPT-5.6: 3 Models, 30x Price Spread, 1 Enterprise Decision
OpenAI just split GPT-5.6 into three models — Sol ($5/$30), Terra ($2.50/$15), and Luna ($1/$6) per million tokens. Sol beats Anthropic's restricted Mythos on TerminalBench. Terra matches GPT-5.5 at half the cost. The release is limited to ~20 organizations under a new U.S. government review process, but the pricing and benchmarks are public. Here's how to classify your workloads across all three tiers and prepare for migration before general availability.
June 26, 2026 · 17 min read