正在加载视频...
视频加载失败
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗
33 条评论

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships! ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

@huggingface does the background noise have effect on it?
@huggingface Oh, could you train OpenAI on this please? 😅

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

@huggingface

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

@huggingface @openwhispr @gabrielste1n

@Scobleizer @huggingface @__cski jfyi

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

@huggingface 💚🚀💫

@huggingface i guess we solved diarization..

@huggingface eight people talking at once and the model is taking attendance

@huggingface Sweet. Love me some Diarrheaization.

@huggingface now this looks pretty interesting

@huggingface woah this is huge!

@huggingface This is awesome

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

@huggingface huge useful

@huggingface Geil. Das ist sau stark.
