Yiwen Gao and colleagues introduce DR-Arena, an automated evaluation framework for Deep Research agents that addresses limitations of static benchmarks through dynamic investigation using real-time web trends. DR-Arena generates structured tasks testing deep reasoning and wide coverage, with an adaptive controller that escalates complexity until capability boundaries emerge. The framework achieves a Spearman correlation of 0.94 with the LMSYS Search Arena leaderboard, representing state-of-the-art alignment with human preferences without manual effort.
No score is assigned. Sources and their independence are shown in the citation chain below.