正在加载视频...

视频加载失败

When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗

1,063,795 次观看 • 18 天前 •via X (Twitter)

33 条评论

NVIDIA AI 的头像
NVIDIA AI18 天前

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

NVIDIA AI 的头像
NVIDIA AI18 天前

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

minime · fireply.ai 的头像
minime · fireply.ai18 天前

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Jon Taylor 的头像
Jon Taylor17 天前

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships!‍ ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

Stephen Turner 🇬🇧🇺🇦 的头像
Stephen Turner 🇬🇧🇺🇦18 天前

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

Rompel 的头像
Rompel18 天前

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

Symbioza2025 | ASA | CLM AI 的头像
Symbioza2025 | ASA | CLM AI18 天前

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

drefrajo 的头像
drefrajo18 天前

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

Knowix 的头像
Knowix18 天前

@huggingface does the background noise have effect on it?

Valery 的头像
Valery17 天前

@huggingface Oh, could you train OpenAI on this please? 😅

Fab 的头像
Fab17 天前

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

Marktechpost AI 的头像
Marktechpost AI17 天前

@huggingface

GloktaCore 的头像
GloktaCore18 天前

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

Aditya Kharbanda 的头像
Aditya Kharbanda17 天前

@huggingface @openwhispr @gabrielste1n

Seiji Satō 的头像
Seiji Satō17 天前

@Scobleizer @huggingface @__cski jfyi

Harsh Mishra 的头像
Harsh Mishra17 天前

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

Alcreon 的头像
Alcreon17 天前

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

Rakesh Gohel 🇨🇦 的头像
Rakesh Gohel 🇨🇦18 天前

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

Emmy 的头像
Emmy17 天前

@huggingface 💚🚀💫

PENIEL 的头像
PENIEL17 天前

@huggingface i guess we solved diarization..

Panther 的头像
Panther18 天前

@huggingface eight people talking at once and the model is taking attendance

George 的头像
George17 天前

@huggingface Sweet. Love me some Diarrheaization.

Bailey Simrell 的头像
Bailey Simrell17 天前

@huggingface now this looks pretty interesting

Niko Storni 🇨🇭 的头像
Niko Storni 🇨🇭17 天前

@huggingface woah this is huge!

Robert Piazza 的头像
Robert Piazza18 天前

@huggingface This is awesome

Subhash Yadav 的头像
Subhash Yadav18 天前

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

Suraj 的头像
Suraj17 天前

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

ThisMightWork 的头像
ThisMightWork17 天前

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

SYNTHLEX 的头像
SYNTHLEX18 天前

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

BoyardE 的头像
BoyardE17 天前

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

Ibesh 的头像
Ibesh18 天前

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

路克0xLUKE777crypt 的头像
路克0xLUKE777crypt17 天前

@huggingface huge useful

Manny Kalavera 的头像
Manny Kalavera17 天前

@huggingface Geil. Das ist sau stark.

相关视频