Terminal-Bench 4.0 leaderboards (September 1) show Fable 5 at 44.5% resolution rate versus GPT-5.6's 37.3%, despite OpenAI claiming GPT-5.6 led on version 2.1. The article questions benchmark reliability: Anthropic-reported Frontier-Bench scores favored Opus 5, GPT-5.6 Sol's claimed result is absent from public leaderboards, and GeneBench is OpenAI's own non-peer-reviewed test. It also examines ARC-AGI-3, ExploitBench, deceptive economics, and selective data reporting.
No score is assigned. Sources and their independence are shown in the citation chain below.