Loading video...

Video Failed to Load

Go Home

When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗

1,063,826 views • 18 days ago •via X (Twitter)

33 Comments

NVIDIA AI's profile picture
NVIDIA AI18 days ago

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

NVIDIA AI's profile picture
NVIDIA AI18 days ago

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

minime · fireply.ai's profile picture
minime · fireply.ai18 days ago

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Jon Taylor's profile picture
Jon Taylor18 days ago

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships!‍ ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

Stephen Turner 🇬🇧🇺🇦's profile picture
Stephen Turner 🇬🇧🇺🇦18 days ago

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

Rompel's profile picture
Rompel18 days ago

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

Symbioza2025 | ASA | CLM AI's profile picture
Symbioza2025 | ASA | CLM AI18 days ago

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

drefrajo's profile picture
drefrajo18 days ago

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

Knowix's profile picture
Knowix18 days ago

@huggingface does the background noise have effect on it?

Valery's profile picture
Valery18 days ago

@huggingface Oh, could you train OpenAI on this please? 😅

Fab's profile picture
Fab18 days ago

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

Marktechpost AI's profile picture
Marktechpost AI17 days ago

@huggingface

GloktaCore's profile picture
GloktaCore18 days ago

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

Aditya Kharbanda's profile picture
Aditya Kharbanda18 days ago

@huggingface @openwhispr @gabrielste1n

Seiji Satō's profile picture
Seiji Satō18 days ago

@Scobleizer @huggingface @__cski jfyi

Harsh Mishra's profile picture
Harsh Mishra18 days ago

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

Alcreon's profile picture
Alcreon18 days ago

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

Rakesh Gohel 🇨🇦's profile picture
Rakesh Gohel 🇨🇦18 days ago

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

Emmy's profile picture
Emmy17 days ago

@huggingface 💚🚀💫

PENIEL's profile picture
PENIEL18 days ago

@huggingface i guess we solved diarization..

Panther's profile picture
Panther18 days ago

@huggingface eight people talking at once and the model is taking attendance

George's profile picture
George18 days ago

@huggingface Sweet. Love me some Diarrheaization.

Bailey Simrell's profile picture
Bailey Simrell18 days ago

@huggingface now this looks pretty interesting

Niko Storni 🇨🇭's profile picture
Niko Storni 🇨🇭18 days ago

@huggingface woah this is huge!

Robert Piazza's profile picture
Robert Piazza18 days ago

@huggingface This is awesome

Subhash Yadav's profile picture
Subhash Yadav18 days ago

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

Suraj's profile picture
Suraj17 days ago

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

ThisMightWork's profile picture
ThisMightWork18 days ago

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

SYNTHLEX's profile picture
SYNTHLEX18 days ago

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

BoyardE's profile picture
BoyardE18 days ago

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

Ibesh's profile picture
Ibesh18 days ago

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

路克0xLUKE777crypt's profile picture
路克0xLUKE777crypt18 days ago

@huggingface huge useful

Manny Kalavera's profile picture
Manny Kalavera18 days ago

@huggingface Geil. Das ist sau stark.

Related Videos