An experiment on GitHub (vyang472/five-bugs) found that 26 agents, from a 4-bit quantized 7B model to Opus 5, all passed the visible test suite on a regex bug while their fixes failed hidden tests. Claude Code and Codex CLI each scored 5/5 on given checkers; both failed the same hidden case, as did Sonnet 5 and Haiku 4.5. Handing one agent the hidden checker fixed it immediately: the test suite, not the model, was the constraint.
No score is assigned. Sources and their independence are shown in the citation chain below.