Researchers from Meta FAIR, University of Oxford and University College London introduced AI Research Preference Models (RPMs), which rank unexecuted research candidates so agents execute only the most promising one. Tested on AIRS-Bench with Qwen3.6-27B, both variants raised average normalized score from 0.684 to 0.711 and 0.729, reaching the baseline's 24-hour score in roughly 15 hours. The team reports new SOTA results on WinoGrande (94.1%) and SVAMP (95.7%). The AIRA-dojo scaffold is open source.
No score is assigned. Sources and their independence are shown in the citation chain below.