← Back to the wire

How to Haircut an AI Benchmark Score

SpeculationBenchmarkSep 30, 2026

On the OfficeQA Pro benchmark built from U.S. Treasury Bulletions, giving models clean text of target pages raised scores by 19 to 29 percentage points, exceeding the 10 to 19 points gained from newer models. The article argues public finance-agent benchmark scores overstate real-world reliability, since clean inputs, preselected documents, and single averaged runs hide ingestion and variance issues that surface when professionals deploy ChatGPT or Claude Cowork on actual deal workflows.

Receipt № 21801 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

ChatGPTModelClaude CoworkModel
Canonical: https://lessuncertain.substack.com/p/how-to-read-and-haircut-ai-benchmark