← Back to the wire

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

AchievementBenchmarkJul 14, 2026

FinResearchBench II introduces a scalable pipeline for generating evaluation rubrics for financial deep research reports without human experts in the final loop. Built from 104 real-world queries, the benchmark synthesizes 14,450 candidate rubrics and retains 2,600 consensus-derived gold rubrics through consistency and distinguishability filters. LLM-based evaluation achieved 98.67% agreement with human experts on unanimous items, enabling differentiated rankings across 10 deep research systems with pass rates from 58.58% to 22.23%.

Receipt № 6751 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Canonical: https://arxiv.org/abs/2607.12252v1