Measuring benchmark optimization in speech recognition
Research found that some ASR models use acoustic cues to optimize for public benchmarks, reproducing reference transcripts even when the audio contradicts them.
Why it matters
Benchmark scores can overstate real-world performance, meaning users might choose a model that is less accurate on new data than its score suggests.
The details
- Some high-scoring models reproduced erroneous reference transcripts 18–30% of the time.
- Models recovered silenced numbers from benchmarks that were not present in the audio.
- Systems adjusted spelling conventions based on which benchmark dataset they identified via acoustics.
Show entities and relationshipsHide entities and relationships
In this article
Products
Companies
Topics
Technologies
Organizations
Key connections
Artificial Analysis owns Real World VoiceEQ
Artificial Analysis introduced the Real World VoiceEQ benchmark suite.
Artificial Analysis owns Open-ASR Leaderboard
Artificial Analysis maintains the Open-ASR Leaderboard.
Artificial Analysis owns Far-field ASR Leaderboard
Artificial Analysis maintains the Far-field ASR Leaderboard.
Cohere owns cohere-transcribe-03-2026
Cohere developed the cohere-transcribe-03-2026 speech model.
NVIDIA owns canary-qwen-2.5b
NVIDIA developed the canary-qwen-2.5b speech recognition model.
IBM owns granite-speech-4.1-2b
IBM developed the granite-speech-4.1-2b speech recognition model.
Show 29 more connectionsShow fewer connections
Microsoft owns Phi-4-multimodal-instruct
Microsoft developed the Phi-4-multimodal-instruct model.
NVIDIA owns parakeet-tdt-0.6b-v2
NVIDIA developed the parakeet-tdt-0.6b-v2 speech recognition model.
Boson AI owns higgs-audio-v3-8b-stt-v2
Boson AI developed the higgs-audio-v3-8b-stt-v2 speech model.
Alibaba owns Qwen3-ASR-0.6B-hf
Alibaba developed the Qwen3-ASR-0.6B-hf speech model.
Mistral AI owns Voxtral-Mini-3B-2507
Mistral AI developed the Voxtral-Mini-3B-2507 speech model.
Moonshot AI owns Kimi-Audio-7B-Instruct
Moonshot AI developed the Kimi-Audio-7B-Instruct model.
OpenAI owns whisper-large-v3
OpenAI developed the whisper-large-v3 speech recognition model.
Moonshine AI owns moonshine-streaming-medium
Moonshine AI developed the moonshine-streaming-medium speech model.
VoxPopuli is built with European Parliament
VoxPopuli benchmark transcripts are derived from European Parliament plenary recordings.
LibriSpeech is built with LibriVox
LibriSpeech dataset is compiled from audiobooks read by LibriVox volunteers.
Artificial Analysis uses VoxPopuli
Artificial Analysis released a cleaned version of VoxPopuli and evaluated models on its transcripts.
Artificial Analysis uses LibriSpeech
Artificial Analysis evaluated speech recognition models against LibriSpeech audio samples.
Open-ASR Leaderboard is related to Automatic Speech Recognition
Open-ASR Leaderboard benchmarks automatic speech recognition systems.
Real World VoiceEQ is related to Automatic Speech Recognition
Real World VoiceEQ evaluates voice AI and speech recognition models.
Far-field ASR Leaderboard is related to Automatic Speech Recognition
Far-field ASR Leaderboard tracks speech recognition performance under far-field conditions.
Artificial Analysis uses Phoneme Error Rate
Artificial Analysis used phoneme error rate probes to measure model transcription fidelity.
Open-ASR Leaderboard uses Word Error Rate
Open-ASR Leaderboard uses word error rate metrics to evaluate speech recognition models.
Artificial Analysis uses Text-to-Speech
Artificial Analysis used text-to-speech voice clones to test model reliance on acoustic cues.
Open-ASR Leaderboard is related to Benchmark Optimization
Open-ASR Leaderboard added a benchmark fitting tab to track benchmark optimization.
Open-ASR Leaderboard uses whisper-large-v3
Open-ASR Leaderboard evaluated whisper-large-v3 on benchmark transcripts and held-out sets.
whisper-large-v3 competes with Voxtral-Mini-3B-2507
whisper-large-v3 and Voxtral-Mini-3B-2507 are evaluated side-by-side as open-source ASR models.
canary-qwen-2.5b competes with parakeet-tdt-0.6b-v2
canary-qwen-2.5b and parakeet-tdt-0.6b-v2 are evaluated side-by-side as speech recognition models.
Phi-4-multimodal-instruct competes with cohere-transcribe-03-2026
Phi-4-multimodal-instruct and cohere-transcribe-03-2026 are evaluated side-by-side.
granite-speech-4.1-2b competes with higgs-audio-v3-8b-stt-v2
granite-speech-4.1-2b and higgs-audio-v3-8b-stt-v2 are evaluated side-by-side.
Qwen3-ASR-0.6B-hf competes with Kimi-Audio-7B-Instruct
Qwen3-ASR-0.6B-hf and Kimi-Audio-7B-Instruct are evaluated side-by-side.
moonshine-streaming-medium competes with whisper-large-v3
moonshine-streaming-medium and whisper-large-v3 are evaluated side-by-side.
OpenAI competes with Mistral AI
OpenAI and Mistral AI compete in developing AI foundation and speech models.
NVIDIA competes with Microsoft
NVIDIA and Microsoft compete in developing speech and multimodal AI models.
Cohere and OpenAI compete in developing AI language and speech models.
Related events
Artificial Analysis Evaluates Benchmark Optimization in Speech Recognition Models
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.