← Back to the wire

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

AchievementBenchmarkSep 4, 2026

On ARC-AGI-3, OpenAI's GPT-6 Astra scored 62.7 percent on ARC Prize's internal harness, surpassing the average human tester's efficiency for the first time. Epoch AI ranks it first overall with 169 points, while Artificial Analysis rates it level with predecessor Sol at 61, behind Claude Fable 5.1's 66. ARC Prize's François Chollet calls the progress "2x faster" than he expected and is moving up his AGI forecast.

Receipt № 17421 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenAICompanyAnthropicCompanyClaude Fable 5ModelEpoch AICompanySolModelArtificial AnalysisCompanyOpus 5ModelARC PrizeCompanyClaude Fable 5.1ModelGPT-6 AstraModelFrançois CholletPerson
Canonical: https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward/