On the OfficeQA Pro benchmark built from U.S. Treasury Bulletions, giving models clean text of target pages raised scores by 19 to 29 percentage points, exceeding the 10 to 19 points gained from newer models. The article argues public finance-agent benchmark scores overstate real-world reliability, since clean inputs, preselected documents, and single averaged runs hide ingestion and variance issues that surface when professionals deploy ChatGPT or Claude Cowork on actual deal workflows.
No score is assigned. Sources and their independence are shown in the citation chain below.