model routingModel Router Buyer's Guide: Buy Failover, Not Judgment
A model router sells two things: mechanical failover, which works, and a learned classifier that picks your model, which the neutral benchmarks say does not. Buy the first, build the second, and plan on 35% savings rather than 60%.
September 11, 2026 · 20 min readagent evaluationOnly 3 of 36 Model Gaps Were Real. Size Your Eval Set.
The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.
September 10, 2026 · 13 min readagent evaluationDeepSeek Won Solo. Gemini Won the Arena. Run Both Rounds.
ERPBench ran 100 identical ERP problems through six model families twice — solo against fixed opponents, then all six competing in one market. DeepSeek won the first, Gemini the second, and the two agreed on 21 of 100 tasks.
September 7, 2026 · 12 min readretrieval evaluationJPMorgan Ranked 62 Retrievers for $800. Pool the Judgments.
JPMorganChase compared 62 retrieval configurations on a production financial-news QA system for about $800, by judging the union of retrieved documents once and reusing 79.6% of the labels. Its top five embedding configurations spanned 0.007 MAP.
September 4, 2026 · 12 min readAI coding agentsClaude Matched the Patch. Qwen Overshot. Score the Scope.
A study of 14,922 agent trajectories across five coding benchmarks finds Claude succeeds by matching the human patch's file scope while Qwen succeeds by exceeding it at every scale. Equal pass rates buy unequal diffs, and your ticket template is the control.
September 3, 2026 · 14 min readAI code reviewBest AI Code Review Tools for a Monorepo: Buy Precision
Four vendors each claim #1 on the same independent benchmark, on four different dates. Greptile holds the highest published precision at 76.2%, but research puts the gap between benchmark scores and real pull requests at 92%.
September 2, 2026 · 14 min readAI vendor shortlistA Demo Vendor Is Perplexity's #3 Source. Rank the Domains.
Two studies published the same day measured what grounds AI software recommendations. Perplexity's third-most-cited domain is a demo-software vendor's marketing blog, 59.8% of its citations sit below Tranco rank 100,000, and 34.7% of numeric citations point at a page that won't open or doesn't contain the number.
September 2, 2026 · 13 min readenterprise searchGlean vs Copilot vs Dust: Buy the Permission Model
Microsoft now bundles Copilot Search free with a $30 Copilot seat and reaches over 100 connectors, so Glean has to beat it on the remainder rather than in the abstract. Dust has 12 connectors and replaces source-system ACLs with its own workspace spaces.
August 28, 2026 · 17 min readStartupBenchAgents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.
StartupBench ran nine frontier agents across 97 workflows taken from products AI startups already sell. A 73.67 rubric average produced an acceptable deliverable on 29.55% of runs — and agents satisfied auxiliary requirements more often than core ones.
August 22, 2026 · 14 min readagent orchestrationAgent Orchestration Platforms: Score Exit, Not Features
Four agent orchestration products shut down, were superseded or repriced in twelve months. A six-criterion scorecard that weights exit cost at 30%, a six-week pilot with a portability test in week three, and the vendor landscape mapped to both.
August 22, 2026 · 18 min readAI agent observabilityBest AI Agent Monitoring: Langfuse, Then a Real Kill Switch
Seven agent monitoring tools priced against the same workload: 50,000 runs a month at 12 steps each. Self-hosted Langfuse wins on cost and audit depth — but only one of these tools sits in the request path where it can actually stop a runaway agent, and that is the part you have to build yourself.
August 11, 2026 · 16 min readvoice agentsEleven Voice Agents, One Bank Call, No Clean Winner
An arXiv benchmark of eleven production voice-agent stacks ran 3,300 simulated calls against a real database. The researchers found task completion collapses on complicated calls, per-call cost varies thirteenfold with no link to quality, and a verification leak hides inside correct refusals.
August 1, 2026 · 13 min read