Felix Feldman extended PubHealthBench into a retrieval-augmented setting, evaluating retrieval and generation choices across 7,929 public health questions. Hybrid retrieval consistently improved recall and ranking quality, while retrieved context substantially increased multiple-choice accuracy, enabling smaller open-weight models to match or outperform larger models without retrieval. A rubric-based LLM-as-a-judge assessed free-form answering, with strongest human agreement on faithfulness and completeness.
No score is assigned. Sources and their independence are shown in the citation chain below.