Artificial Analysis evaluated 11 open-source ASR models and found that several top-scoring systems reproduced benchmark transcripts even when audio contradicted them. Its research introduces three tests to quantify benchmark optimization ("benchmaxxing") in speech recognition. The methodology flagged potential reference errors in 40% of analyzed VoxPopuli test clips, affecting roughly 3% of reference words. Models exhibiting benchmark-optimized behavior reproduced erroneous transcripts 18–30% of the time, and lowest-WER models were most likely to reproduce errors.
No score is assigned. Sources and their independence are shown in the citation chain below.