Haize LabsBeacon Bought Haize Labs. Who Red-Teams Your Agents Now?
Beacon Software acquired AI red-teaming firm Haize Labs to run its internal AI platform, and the announcement says nothing about existing customers. Here is what to secure before renewal.
September 19, 2026 · 11 min readMCPUNICEF Tested a Generic MCP Server. It Lost to No Tools.
On UNICEF's own statistics, a generic SDMX MCP server scored 0.074 at returning the right figure for the right year, below the 0.147 of no tools. A purpose-built server scored 0.990. As the UN opens 44 million data points to agents, the lesson for enterprise data teams is to fund the resolver, not the connector.
September 18, 2026 · 14 min readReal-SWEFable Cost $18 a Fix on Private Code. Flash Cost $8.
On Specific Labs' Real-SWE benchmark of licensed private codebases, Fable 5.1 resolved the most tasks but cost about $17.94 per fix to Gemini 3.8 Flash's $8.01 — and each model fails in its own way, so your review gate should follow the agent you pick.
September 13, 2026 · 14 min readMCP150 MCP Servers Never Started. No Benchmark Shows That.
An unrepaired probability sample of 400 MCP servers found only 48.8% complete an initialize handshake, and 37.5% never start at all. Meanwhile 68.8% of BFCL v4's extracted tool definitions and 85.6% of UltraTool's are exact repeats, against 0.4% for real MCP tools.
September 12, 2026 · 11 min readagent evaluationOnly 3 of 36 Model Gaps Were Real. Size Your Eval Set.
The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.
September 10, 2026 · 13 min readagent evaluationDeepSeek Won Solo. Gemini Won the Arena. Run Both Rounds.
ERPBench ran 100 identical ERP problems through six model families twice — solo against fixed opponents, then all six competing in one market. DeepSeek won the first, Gemini the second, and the two agreed on 21 of 100 tasks.
September 7, 2026 · 12 min readAI coding agents221 Green Patches Failed Review. Put Your Rules in Context.
SWE-Gate scored review-constraint compliance separately from functional tests across 303 repository-level repair tasks. Of 644 agent patches that passed the tests, 221 violated constraints taken from the original pull request reviews.
September 6, 2026 · 13 min readcode hallucination12 Models Wrote a Fake Crate. None Refused a Real One.
A new benchmark handed twelve open-weight models 270 impossible coding tasks. They fabricated confident, compiling code on 60% and refused 27% — while wrongly refusing 0.0% of 91 matched solvable controls. That zero is why an unsatisfiable arm belongs in your eval suite.
September 5, 2026 · 13 min readAI coding agentsClaude Matched the Patch. Qwen Overshot. Score the Scope.
A study of 14,922 agent trajectories across five coding benchmarks finds Claude succeeds by matching the human patch's file scope while Qwen succeeds by exceeding it at every scale. Equal pass rates buy unequal diffs, and your ticket template is the control.
September 3, 2026 · 14 min readagent memoryAgent Memory Cost 14 Points at Best. Test With It Off.
MemTrapBench tested five agent-memory frameworks against a no-memory baseline on multi-turn business dialogues. All five scored worse — 85.16% down to 71.17% at best on Gemini-3-Flash — because correct, relevant memories anchor the model to the previous task's framing.
August 23, 2026 · 12 min readStartupBenchAgents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.
StartupBench ran nine frontier agents across 97 workflows taken from products AI startups already sell. A 73.67 rubric average produced an acceptable deliverable on 29.55% of runs — and agents satisfied auxiliary requirements more often than core ones.
August 22, 2026 · 14 min readQwen3.8-MaxAlibaba's Agent Coded 16 Days. A Human Wrote 13 Commits.
Alibaba says Qwen3.8-Max coded unattended for 16 days, and unlike every rival long-horizon claim it shipped the whole commit trace on public GitHub. Open it and the 16 days become 'more than ten' in Alibaba's own words, 13 of 648 commits turn out to be a human's, and one person still holds the merge keys.
August 5, 2026 · 10 min read