Measuring benchmark optimization in speech recognition
Benchmark scores can overstate real-world performance, meaning users might choose a model that is less accurate on new data than its score suggests.
- Evaluated 11 open-source automatic speech recognition models using consensus disagreement, masked entity retrieval, and orthographic switching probes
- Observed that six of 11 models reproduced erroneous reference transcripts from VoxPopuli rather than faithfully transcribing the actual spoken audio
- Demonstrated that top-scoring models on LibriSpeech reproduced silenced numbers in roughly 30% to 40% of test examples despite the audio being removed
- Found models rely on acoustic cues to select benchmark-specific spelling conventions, exceeding random choice with up to 90% switch accuracy
- Introduced a Benchmark fitting tab to the Open-ASR Leaderboard to identify models over-optimized for public benchmarks