← Back to the wire

Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

AchievementBenchmarkJul 10, 2026

Ethan Leung and colleagues found GPT-5-mini attained the strongest source-relevance pass-class F1 at 0.908 on an adversarial benchmark evaluating 8 LLM judges for citation quality. Across 1,248 rubric decisions, cheaper judges remain competitive; on factual support, judges are statistically indistinguishable. Scalar F1 obscures directional biases in pass-rate drift and false positive and negative rates that reinforcement learning loops reinforce, making calibration a prerequisite for reward signals.

Receipt № 4521 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Ethan LeungPersonElias LumerPersonCorey FeldPersonAustin HuberPersonVamse Kumar SubbiahPersonKevin PaulPerson
Canonical: https://arxiv.org/abs/2607.08700