
Announcing the Speech Agent Arena: Compare Speech agents in real world conversations | Artificial Analysis
Artificial Analysis launched the Speech Agent Arena to evaluate Speech to Speech models using human preference and task success rates in real-world scenarios.
Why it matters
It helps users distinguish between voice agents that sound natural and those that are most reliable at completing actual tasks.
The details
- Gemini 3.1 Flash Live Preview - Minimal leads in conversational preference Elo.
- SpaceXAI Grok Voice Think Fast 2.0 High leads in task success rate.
- Lower Time to First Audio generally increases conversational preference.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.