← Back to the wire

Adding Benchmaxxer Repellant to the Open ASR Leaderboard

AnnouncementBenchmarkMay 6, 2026

Appen Inc. and DataoceanAI have provided private English ASR datasets to Hugging Face's Open ASR Leaderboard to prevent benchmark-specific optimization. The datasets cover scripted and conversational speech across multiple accents. The leaderboard's Average WER remains computed on public datasets only, with an optional toggle to include private data. Contributors Eric Bezzam, Steven Zheng, and others led the effort to improve benchmark trustworthiness against test-set contamination.

Receipt № 4991 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

Appen Inc.CompanyDataoceanAICompanyEric BezzamPersonEustache Le BihanPersonSergio BruccoleriPersonJeanine Sinanan-SinghPersonCasey FordPersonGuanbo WangPersonYukai HuangPersonKe LiPersonYufeng HaoPersonLiao XiaolingPersonSteven ZhengPersonHugging FaceCompany
Canonical: https://huggingface.co/blog/open-asr-leaderboard-private-data