Real-SWEFable Cost $18 a Fix on Private Code. Flash Cost $8.
On Specific Labs' Real-SWE benchmark of licensed private codebases, Fable 5.1 resolved the most tasks but cost about $17.94 per fix to Gemini 3.8 Flash's $8.01 — and each model fails in its own way, so your review gate should follow the agent you pick.
September 13, 2026 · 14 min readAI coding agents221 Green Patches Failed Review. Put Your Rules in Context.
SWE-Gate scored review-constraint compliance separately from functional tests across 303 repository-level repair tasks. Of 644 agent patches that passed the tests, 221 violated constraints taken from the original pull request reviews.
September 6, 2026 · 13 min readAI coding agentsZalando Auto-Approves a Third of PRs. Agents Made Them Bigger.
Zalando published 2.5 years of agentic engineering data across 250+ teams. The 20-40% pull request lead-time win came from a bot that auto-approves 33% of PRs without a human — while PR sizes climbed into the 1k-2k line buckets and per-commit cyclomatic complexity showed inflection points exactly where coding agents entered.
August 17, 2026 · 13 min read