OpenAI's GPT-6 Astra scored 80% on Epoch AI's Furniture Assembly Benchmark, which tests whether models can identify deliberate errors in IKEA assembly photos. Claude Opus 4.5 managed 28% in November 2025; Claude Fable 5.1 reaches 70% and Claude Opus 5 61%. Open-weight models like Kimi K3 trail by at least seven months. Researchers say the tech remains too slow for real-time help but could aid car or appliance repairs.
No score is assigned. Sources and their independence are shown in the citation chain below.