Anthropic published a paper showing automated AI researchers improved performance on all 10 alignment benchmarks without degrading overall model performance. Led by Anthropic fellow Chen Yueh-Han, the system searches literature, proposes methods, and trains models in 30-minute iterations, keeping effective methods. The paper reports the best automated method beats experienced human proposals within six hours, at roughly $4 per hour versus $150 for human researchers. Limitations include dependence on benchmark quality.
No score is assigned. Sources and their independence are shown in the citation chain below.