Video yükleniyor...
Video Yüklenemedi
Introducing Jarvis Bench v0.5, Voice Arena's conversational agent benchmark. We've been obsessed with one question at VoiceArena: why do voice agent demos sound incredible, benchmarks say models are near-perfect, and yet you probably didn't have a single real conversation with a voice agent in the last 24 hours? Here's... show more
568,930 görüntüleme • 10 gün önce •via X (Twitter)
12 Yorum

The first board is Naturalness: how much each call sounds like a real human conversation. The human agent sits at 1244 Elo. Every model is a wide gap below the human: › gpt-realtime 984 (-260) › Gemini 3.1 Flash Live 970 (-274) › GPT Live 1 967 (-277) › Grok Voice 920 (-324) › Inworld cascaded 915 (-329) The best model wins just 18% of blind calls against a human. No model is above 18%. Sounding human is where the gap is still widest.

Read every pair head to head and the same thing holds. The human agent is favoured against every single model: 82% against gpt-realtime, and up to 87% against Grok Voice and Inworld. The models sit within 60-40 of each other, so the race between them is close. The race against a person is not: none has crossed 50% yet.

Here's one of those calls: gpt-realtime, our top model on Humanness, next to a real caller, on a gift-budget question. Same conversation, side by side. Play it and see how far 18% actually sounds.

On Task Completion, though, it's a different story, the gap is much closer. Votes are still coming in and we will keep updating the leaderboard. This is a work in progress. Tag the model labs you would like to see added to this benchmark.

@voicearena_ai yes, I suspect voice benchmarks are going to become much more multidimensional than text benchmarks.

@voicearena_ai most systems are trained to sound like polished audiobook narrators. real human dialogue is messy, with false starts and micro-interruptions. until models learn acoustic imperfection, that naturalness gap stays wide open

@voicearena_ai This is great and much needed! Congratulations to the entire team 🙌🏼

@voicearena_ai Demos and clean benchmarks fail the same way: they never see barge-in on dirty audio mid-tool-call. Blind human ratings on real conversations is the right bar. Task completion without sounding robotic is the harder half.

@voicearena_ai I think real conversations are a much better test than demos.

@voicearena_ai Real conversations reveal what benchmarks miss

@voicearena_ai blind votes are closer to the real test than a transcript score. the failure mode is usually turn-taking and recovery after a bad ASR guess, not the first answer.

@voicearena_ai @shobhitbanga sounds like your 'bench' has some ghostly AI 'users' 🤖😅
