
12 Models Wrote a Fake Crate. None Refused a Real One.
A new benchmark handed twelve open-weight models 270 impossible coding tasks. They fabricated confident, compiling code on 60% and refused 27% — while wrongly refusing 0.0% of 91 matched solvable controls. That zero is why an unsatisfiable arm belongs in your eval suite.
September 5, 2026 · 13 min read