OpenAI's GPT-6 Astra ranked first on ulam.ai's ErdosBench, scoring 3.23 and solving 106 of 226 open math problems, 43 completely. Benchmark developer Przemek Chojecki described it as a 5%-10% gain over Sol, which solved 78 problems. OpenAI chief scientist Jakub Pachocki said the company deliberately did not prioritize math, focusing instead on recursive self-improvement and automated alignment research. Compared to Sol, Astra showed stronger scientific writing and fewer overblown claims.
No score is assigned. Sources and their independence are shown in the citation chain below.