Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗

1,063,826 Aufrufe • vor 18 Tagen •via X (Twitter)

33 Kommentare

Profilbild von NVIDIA AI
NVIDIA AIvor 18 Tagen

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

Profilbild von NVIDIA AI
NVIDIA AIvor 18 Tagen

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

Profilbild von minime · fireply.ai
minime · fireply.aivor 18 Tagen

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Profilbild von Jon Taylor
Jon Taylorvor 18 Tagen

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships!‍ ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

Profilbild von Stephen Turner 🇬🇧🇺🇦
Stephen Turner 🇬🇧🇺🇦vor 18 Tagen

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

Profilbild von Rompel
Rompelvor 18 Tagen

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

Profilbild von Symbioza2025 | ASA | CLM AI
Symbioza2025 | ASA | CLM AIvor 18 Tagen

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

Profilbild von drefrajo
drefrajovor 18 Tagen

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

Profilbild von Knowix
Knowixvor 18 Tagen

@huggingface does the background noise have effect on it?

Profilbild von Valery
Valeryvor 18 Tagen

@huggingface Oh, could you train OpenAI on this please? 😅

Profilbild von Fab
Fabvor 18 Tagen

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

Profilbild von Marktechpost AI
Marktechpost AIvor 17 Tagen

@huggingface

Profilbild von GloktaCore
GloktaCorevor 18 Tagen

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

Profilbild von Aditya Kharbanda
Aditya Kharbandavor 18 Tagen

@huggingface @openwhispr @gabrielste1n

Profilbild von Seiji Satō
Seiji Satōvor 18 Tagen

@Scobleizer @huggingface @__cski jfyi

Profilbild von Harsh Mishra
Harsh Mishravor 18 Tagen

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

Profilbild von Alcreon
Alcreonvor 18 Tagen

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

Profilbild von Rakesh Gohel 🇨🇦
Rakesh Gohel 🇨🇦vor 18 Tagen

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

Profilbild von Emmy
Emmyvor 17 Tagen

@huggingface 💚🚀💫

Profilbild von PENIEL
PENIELvor 18 Tagen

@huggingface i guess we solved diarization..

Profilbild von Panther
Panthervor 18 Tagen

@huggingface eight people talking at once and the model is taking attendance

Profilbild von George
Georgevor 18 Tagen

@huggingface Sweet. Love me some Diarrheaization.

Profilbild von Bailey Simrell
Bailey Simrellvor 18 Tagen

@huggingface now this looks pretty interesting

Profilbild von Niko Storni 🇨🇭
Niko Storni 🇨🇭vor 18 Tagen

@huggingface woah this is huge!

Profilbild von Robert Piazza
Robert Piazzavor 18 Tagen

@huggingface This is awesome

Profilbild von Subhash Yadav
Subhash Yadavvor 18 Tagen

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

Profilbild von Suraj
Surajvor 17 Tagen

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

Profilbild von ThisMightWork
ThisMightWorkvor 18 Tagen

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

Profilbild von SYNTHLEX
SYNTHLEXvor 18 Tagen

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

Profilbild von BoyardE
BoyardEvor 18 Tagen

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

Profilbild von Ibesh
Ibeshvor 18 Tagen

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

Profilbild von 路克0xLUKE777crypt
路克0xLUKE777cryptvor 18 Tagen

@huggingface huge useful

Profilbild von Manny Kalavera
Manny Kalaveravor 18 Tagen

@huggingface Geil. Das ist sau stark.

Ähnliche Videos