Measuring benchmark optimization in speech recognition
| Source: Hugging Face Blog
Tags: ASR, speech-recognition, benchmarking, Hume-AI, Open-ASR-Leaderboard, VoxPopuli, benchmark-gaming
Hume AI researchers tested 11 open-source ASR models and found that several reproduce known benchmark transcripts even when the audio is silenced or contradicted — systematic benchmark gaming that inflates leaderboard scores beyond real-world accuracy.
Details
Researchers from Hume AI, published on Hugging Face, document a measurable benchmaxxing phenomenon in automatic speech recognition: models that top public leaderboards appear to score well not because they transcribe speech more accurately, but because they have learned to reproduce expected outputs for known benchmark recordings. The team designed three tests to probe this. First, a consensus disagreement probe on VoxPopuli (which has known transcription errors) checks whether models transcribe what the audio actually says or reproduce the benchmark's incorrect reference. Second, a word-silencing test removes target words from audio to see if models still produce them. Third, acoustic-cue detection tests whether models identify which benchmark they are being evaluated on from subtle audio features. Evaluating 11 widely-used open-source ASR models on VoxPopuli English and LibriSpeech, several high-ranking systems showed all three patterns. To counter benchmark gaming, the authors introduced held-out test sets in three public benchmarks: Real World VoiceEQ, the Open-ASR Leaderboard, and the Far-field ASR Leaderboard. These held-out sets were not used during training and expose models that have overfit to known test distributions. The practical implication for teams deploying speech AI is direct: a model with near-human leaderboard scores may still fail unpredictably in production when audio conditions diverge from benchmark distributions. Evaluation on held-out or domain-specific sets is more predictive of real-world performance than public rankings.