Artificial Analysis released Intelligence Index v4.2, an interim update adding AA-Briefcase for agentic knowledge work and Surge's GDP.pdf evaluation of long-context document reasoning across 4,592 pages. GPQA Diamond was removed as saturated. Private held-out test sets now account for 40% of Index weighting, up from 20% in v4.1. Anthropic's Claude Fable 5.1 leads the Index, followed by OpenAI's GPT-6 Astra.
No score is assigned. Sources and their independence are shown in the citation chain below.