Hugging Face launched Community Evals, a decentralized evaluation system where benchmark datasets host leaderboards and models store their own eval scores in .eval_results/*.yaml files. Ben Burtenshaw, Nathan Habib, Bertrand Chevrier, Merve, Daniel van Strien, Niels Rogge, and Julien Chaumond announced the feature in beta. Any user can submit evaluation results for models via pull requests, with verified badges confirming reproducibility. MMLU-Pro, GPQA, and HLE benchmarks are already live, with plans to expand. All scores are accessible via Hub APIs.
No score is assigned. Sources and their independence are shown in the citation chain below.