← Back to the wire

GitHub - vyang472/five-bugs: A 10-minute smoke test for AI coding agents. Five seeded Python bugs, one checker the agent sees and one it never does — five agents across two labs and three model tiers all pass the first and fail the same case in the second.

AchievementBenchmarkSep 15, 2026

An experiment on GitHub (vyang472/five-bugs) found that 26 agents, from a 4-bit quantized 7B model to Opus 5, all passed the visible test suite on a regex bug while their fixes failed hidden tests. Claude Code and Codex CLI each scored 5/5 on given checkers; both failed the same hidden case, as did Sonnet 5 and Haiku 4.5. Handing one agent the hidden checker fixed it immediately: the test suite, not the model, was the constraint.

Receipt № 19251 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01highPRIMARY
GitHubCompanyClaude CodeModelSonnet 5ModelHaiku 4.5ModelOpus 5ModelCodex CLICompany
Canonical: https://github.com/vyang472/five-bugs