Loading video...
Video Failed to Load
Introducing Diarization Bench, Voice Arena's benchmark for who spoke when. The metric we measure is diarization error rate, or DER. It measures how much of a conversation a model attaches to the wrong person. Every second it misses, invents, or gives to another speaker counts against it. Here's what... show more
628,174 views • 2 days ago •via X (Twitter)
7 Comments

The leaderboard at a 0 ms collar @nvidia Nemotron 3 leads at 14.72%. BUTFIT Speech’s DiariZen v2 follows at 19.34, PyannoteAI’s Precision-3 at 20.56 and Precision-2 at 23.40. Deepgram’s Nova-3 is last at 67.13 Open weights and commercial APIs, same audio, no speaker count given to anyone.

Diarization error rate, or DER: It measures how much of a conversation a model attaches to the wrong person. Every second it misses, invents, or confuses with another speaker counts against it and increases the error. 0% DER would be a transcript where nobody is ever misquoted and every speaker label is correct. Then there is the collar. When one person stops and the next starts, the exact moment of the handover is blurry, so most benchmarks agree to ignore a small window around every speaker change, and nothing counts as a mistake inside it. Make that window wider and the mistakes disappear, which is why the same model can report two very different scores. Voice Arena measures DER at three collars: 0, 100 and 250 ms. The corpus: about 22 hours and 280+ speakers. Treated rooms, open offices, courtyards and busy public spaces, two to five people in a room and two to eight on a call, with overlap going past 30% to stress test the models. 12 systems, open weights and commercial APIs, one protocol, with a single goal: to test these models in the conditions they are actually put in.

The same system can report several different numbers, depending on one setting: the collar. At the 250 ms collar, Nemotron 3 reads 4.29%. Score the same system on the same audio at 0 ms collar, and it reads 14.72%. That is 3.4x the error, from one convention.

Where it actually breaks: overlap, and how many people are in the room. Nemotron 3 goes from 12.0% to 17.9% once overlap passes 30%, and from 10.3% on two speakers to 22.2% on five or more. AssemblyAI’s Universal-3.5 Pro goes from 26.5% to 58.3%. A crowded, overlapping room costs some systems far more than others. Across the board, the pattern holds: Nemotron 3 and DiariZen stay under 26% even in the hardest rooms, the Pyannote family sits in the middle, and everything down from there loses more than half the conversation once the room fills up. Diarization is not solved, and it is still a long way from it. Now live on

@voicearena_ai really meaningful benchmark. and looks like I need to migrate to NVIDIA Nemotron 3

@voicearena_ai Awesome breakdown 👊

@voicearena_ai Nemotron 3 is goated ..
