Ethan Leung and colleagues found GPT-5-mini attained the strongest source-relevance pass-class F1 at 0.908 on an adversarial benchmark evaluating 8 LLM judges for citation quality. Across 1,248 rubric decisions, cheaper judges remain competitive; on factual support, judges are statistically indistinguishable. Scalar F1 obscures directional biases in pass-rate drift and false positive and negative rates that reinforcement learning loops reinforce, making calibration a prerequisite for reward signals.
No score is assigned. Sources and their independence are shown in the citation chain below.