Loading video...

Video Failed to Load

Go Home

Introducing Diarization Bench, Voice Arena's benchmark for who spoke when. The metric we measure is diarization error rate, or DER. It measures how much of a conversation a model attaches to the wrong person. Every second it misses, invents, or gives to another speaker counts against it. Here's what...

628,174 views • 2 days ago •via X (Twitter)

7 Comments

Shobhit Banga's profile picture
Shobhit Banga2 days ago

The leaderboard at a 0 ms collar @nvidia Nemotron 3 leads at 14.72%. BUTFIT Speech’s DiariZen v2 follows at 19.34, PyannoteAI’s Precision-3 at 20.56 and Precision-2 at 23.40. Deepgram’s Nova-3 is last at 67.13 Open weights and commercial APIs, same audio, no speaker count given to anyone.

Shobhit Banga's profile picture
Shobhit Banga2 days ago

Diarization error rate, or DER: It measures how much of a conversation a model attaches to the wrong person. Every second it misses, invents, or confuses with another speaker counts against it and increases the error. 0% DER would be a transcript where nobody is ever misquoted and every speaker label is correct. Then there is the collar. When one person stops and the next starts, the exact moment of the handover is blurry, so most benchmarks agree to ignore a small window around every speaker change, and nothing counts as a mistake inside it. Make that window wider and the mistakes disappear, which is why the same model can report two very different scores. Voice Arena measures DER at three collars: 0, 100 and 250 ms. The corpus: about 22 hours and 280+ speakers. Treated rooms, open offices, courtyards and busy public spaces, two to five people in a room and two to eight on a call, with overlap going past 30% to stress test the models. 12 systems, open weights and commercial APIs, one protocol, with a single goal: to test these models in the conditions they are actually put in.

Shobhit Banga's profile picture
Shobhit Banga2 days ago

The same system can report several different numbers, depending on one setting: the collar. At the 250 ms collar, Nemotron 3 reads 4.29%. Score the same system on the same audio at 0 ms collar, and it reads 14.72%. That is 3.4x the error, from one convention.

Shobhit Banga's profile picture
Shobhit Banga2 days ago

Where it actually breaks: overlap, and how many people are in the room. Nemotron 3 goes from 12.0% to 17.9% once overlap passes 30%, and from 10.3% on two speakers to 22.2% on five or more. AssemblyAI’s Universal-3.5 Pro goes from 26.5% to 58.3%. A crowded, overlapping room costs some systems far more than others. Across the board, the pattern holds: Nemotron 3 and DiariZen stay under 26% even in the hardest rooms, the Pyannote family sits in the middle, and everything down from there loses more than half the conversation once the room fills up. Diarization is not solved, and it is still a long way from it. Now live on

Rohan Paul's profile picture
Rohan Paul1 day ago

@voicearena_ai really meaningful benchmark. and looks like I need to migrate to NVIDIA Nemotron 3

NVIDIA AI's profile picture
NVIDIA AI1 day ago

@voicearena_ai Awesome breakdown 👊

AshutoshShrivastava's profile picture
AshutoshShrivastava1 day ago

@voicearena_ai Nemotron 3 is goated ..

Related Videos

Today we're announcing $280M in Series B funding at a $2B valuation, led by our long-time investor and partner, Menlo Ventures. When we announced our Series A last May, most conversations I had about voice began with someone explaining why they doubted it. I don't have those conversations anymore. People tell me instead how much time they save talking instead of typing and what they want us to build next. In a little over a year, voice has gone from something people were questioning to something they rely on, and it happened faster than we expected. That shift is why we've raised a new round of funding. This funding represents a deeper investment in our products, lab, models, and our team. Alongside the funding, we're announcing a preview of our first proprietary speech model, Canto. Canto is a 2B parameter speech model trained for the places people actually talk: loud rooms, windy streets, a toddler in the background, a second language mixed into the first. Models like this will change how we all interact with devices, and Canto is the first in a rapid line of them. Each one will be larger and more capable than the last, and each should improve Flow's dictation in a way you can feel the day it ships. Larger speech models also open the door to the vision we dreamed up 5 years ago: using your voice as the primary way you interact with devices. That still has to be invented. A voice interface people rely on all day, in every app, doesn't exist anywhere yet. It's time for that to change. To Sahaj Garg, our team, and every person who gave Wispr Flow a chance - this one's for you.

Tanay Kothari

615,867 views • 1 month ago

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 views • 4 months ago

The antisemitism that was festering under the surface for decades, has exploded. And for those of us who stand against it, we are not the popular kids. The so called “cool kids” are wearing a keffiyeh, calling for an intifada and demanding a Palestine between the river and the sea even if they can’t tell you which river or which sea. Keffiyeh Karens. That’s the mainstream now. Standing up against jihadism is now the counter-culture. Standing up for western democratic values against terror and extremism: that is now the counter-culture. Using facts and reason and knowing history: that is now the counter-culture. People! WE are the counterculture. And we are doing it. The Jewish people have found our voice because standing up for what we believe in, even when it’s not popular, is part of who we are. Our tiny people haven’t survived for millennia and outlived every empire that tried to destroy us by always being the cool kids. In fact, we almost never are. We know when it’s time to swim against the tide because we know this playbook. We have been here before. And our message to the world is as follows: the hatred that targets is targeting you too. I have a question for you, keffiyeh Karen from campus: where would you prefer to be a woman? in Iran? In Afghanistan? or in Israel? We cannot take the freedoms that we enjoy here in the US in Israel and in the west for granted, and right now even saying just that is swimming against the tide – it is being the counterculture. This is not a war of Israel’s choosing but we cannot avoid it either. They might call us the chosen people: in fact we are the people with no choice. But the entire western world needs to wake up to the fact that it has no choice either: that what is at stake is the survival, not just of the Jewish people, but of Western civilization, democracy and freedom itself. And where we do have a choice is in how we do it: and we choose to speak the truth, even when the truth isn’t popular. To speak up, even when we are shouted down and to be proud even when they tell us we better hide. Thank you so much, am Israel chai

Noa Tishby

349,873 views • 2 years ago

YOKO ONO: ONOCHORD, VENICE, 2004 Yoko: The world is divided in two industries. One is the War Industry and the other is the Peace Industry. The people in the War Industry are totally together. They don't have to talk to each other, even. They know exactly what they want to do. They want to go out there, kill and make money. But the people in the Peace Industry, which are us - we are so idealistic that each one of us criticises the other Peace Person in the Peace Industry. And we are always just arguing and we are wasting our energies doing that. So let's just forgive each other and see that we are in the Peace Industry and that's all that counts. Even if you are not marching for peace, just be yourself, being a florist, being a merchant, being a talior, anything. That way you're contributing to the Peace Industry. People are just concentrating on fear, confusion and anger. And therefore just for a moment, I'd like us to think about Love. In a very magical, straight way, John and I met in London and from then on we stood for Peace and Love. And when I do this kind of event. Well it is... I was inspired to do it, but I still think that I'm still with John in spirit. John and I created the country called Nutopia. Not Utopia, because there was Utopia as a concept already. And we wanted to create a new concept, so we just added N on it - Nutopia - and as a country. Well, that is the concept of a country. And we all are citizens of that country. And in my apartment in the Dakota Building, we put a little plaque on the back door, the kitchen door. It says 'Nutopian Embassy' and even now we have that. (laughs). Nutopia exists in our minds. And because of that, some people want to rebel against it. The reason some want to rebel against it is a good proof that it exists. I think that it was a terrible thing that happened in Chechnya. But we have to still keep our hopes up. And instead of giving up, we have to keep on sending the message of Love to each other. You say that I am the Ambassador of Peace. We are all Ambassadors of Peace. You are too. Everybody in this room are Ambassadors of Peace. Just the fact that we are not participating in War. The fact that we are here, and we are what we are, means that we are in the Peace Industry. All of us. John and I used to say that our apartment in the Dakota is a conceptual monastry, just for the two of us. And when we go out of the Dakota, we get so many people communicating with us, so it's very important that we had silence and quietness. And my apartment is a very small space compared to the world. And I need that for my peace of mind. You should be kind to each other. You should come together, hug each other, love each other, express our love to each other and we should make it work. We should finally create a world that is a totally an Earth for Us. So let's do it. Yoko Ono, OpenAsia Press Conference, whilst exhibiting Onochord, 2004 by Yoko Ono (Nutopia) at the Venice Biennale: OpenAsia 2004, Lido Di Venezia, Venice, Italy, 9 September 2004.

Yoko Ono

35,208 views • 2 years ago