
Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.
September 10, 2026 · 13 min read