Opus 5.5 and GPT-6 Astra claim 72.6% and 81.8% on OSWorld 2.0, respectively, but the article argues that self-contained benchmarks remain far simpler than real deployment environments. It contends that open-ended data is the bottleneck for autonomous agents: current trajectory datasets use scripted, single-objective tasks on sterile desktops, missing signals for triaging interruptions, carrying context forward, and navigating multi-threaded, evolving work.
No score is assigned. Sources and their independence are shown in the citation chain below.