← Back to the wire

Community Evals: Because we're done trusting black-box leaderboards over the community

AnnouncementProductFeb 4, 2026

Hugging Face launched Community Evals, a decentralized evaluation system where benchmark datasets host leaderboards and models store their own eval scores in .eval_results/*.yaml files. Ben Burtenshaw, Nathan Habib, Bertrand Chevrier, Merve, Daniel van Strien, Niels Rogge, and Julien Chaumond announced the feature in beta. Any user can submit evaluation results for models via pull requests, with verified badges confirming reproducibility. MMLU-Pro, GPQA, and HLE benchmarks are already live, with plans to expand. All scores are accessible via Hub APIs.

Receipt № 5631 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01highPRIMARY
Ben BurtenshawPersonMervePersonBertrand ChevrierPersonJulien ChaumondPersonNiels RoggePersonNathan HabibPersonDaniel van StrienPersonHugging FaceCompany
Canonical: https://huggingface.co/blog/community-evals