← Back to the wire

Why AI Benchmarks Are Total BS (And How OpenAI and Anthropic Use Them to Trick You)

SpeculationBenchmarkSep 12, 2026

Terminal-Bench 4.0 leaderboards (September 1) show Fable 5 at 44.5% resolution rate versus GPT-5.6's 37.3%, despite OpenAI claiming GPT-5.6 led on version 2.1. The article questions benchmark reliability: Anthropic-reported Frontier-Bench scores favored Opus 5, GPT-5.6 Sol's claimed result is absent from public leaderboards, and GeneBench is OpenAI's own non-peer-reviewed test. It also examines ARC-AGI-3, ExploitBench, deceptive economics, and selective data reporting.

Receipt № 18931 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenAICompanyAnthropicCompanyGPT-5.6ModelOpus 5ModelFable 5.1Model
Canonical: https://www.pcmag.com/opinions/why-ai-benchmarks-are-total-bs-and-how-openai-and-anthropic-use-them-to