Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗

1,063,826 görüntüleme • 18 gün önce •via X (Twitter)

33 Yorum

NVIDIA AI profil fotoğrafı
NVIDIA AI18 gün önce

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

NVIDIA AI profil fotoğrafı
NVIDIA AI18 gün önce

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

minime · fireply.ai profil fotoğrafı
minime · fireply.ai18 gün önce

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Jon Taylor profil fotoğrafı
Jon Taylor18 gün önce

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships!‍ ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

Stephen Turner 🇬🇧🇺🇦 profil fotoğrafı
Stephen Turner 🇬🇧🇺🇦18 gün önce

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

Rompel profil fotoğrafı
Rompel18 gün önce

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

Symbioza2025 | ASA | CLM AI profil fotoğrafı
Symbioza2025 | ASA | CLM AI18 gün önce

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

drefrajo profil fotoğrafı
drefrajo18 gün önce

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

Knowix profil fotoğrafı
Knowix18 gün önce

@huggingface does the background noise have effect on it?

Valery profil fotoğrafı
Valery18 gün önce

@huggingface Oh, could you train OpenAI on this please? 😅

Fab profil fotoğrafı
Fab18 gün önce

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

Marktechpost AI profil fotoğrafı
Marktechpost AI17 gün önce

@huggingface

GloktaCore profil fotoğrafı
GloktaCore18 gün önce

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

Aditya Kharbanda profil fotoğrafı
Aditya Kharbanda18 gün önce

@huggingface @openwhispr @gabrielste1n

Seiji Satō profil fotoğrafı
Seiji Satō18 gün önce

@Scobleizer @huggingface @__cski jfyi

Harsh Mishra profil fotoğrafı
Harsh Mishra18 gün önce

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

Alcreon profil fotoğrafı
Alcreon18 gün önce

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

Rakesh Gohel 🇨🇦 profil fotoğrafı
Rakesh Gohel 🇨🇦18 gün önce

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

Emmy profil fotoğrafı
Emmy17 gün önce

@huggingface 💚🚀💫

PENIEL profil fotoğrafı
PENIEL18 gün önce

@huggingface i guess we solved diarization..

Panther profil fotoğrafı
Panther18 gün önce

@huggingface eight people talking at once and the model is taking attendance

George profil fotoğrafı
George18 gün önce

@huggingface Sweet. Love me some Diarrheaization.

Bailey Simrell profil fotoğrafı
Bailey Simrell18 gün önce

@huggingface now this looks pretty interesting

Niko Storni 🇨🇭 profil fotoğrafı
Niko Storni 🇨🇭18 gün önce

@huggingface woah this is huge!

Robert Piazza profil fotoğrafı
Robert Piazza18 gün önce

@huggingface This is awesome

Subhash Yadav profil fotoğrafı
Subhash Yadav18 gün önce

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

Suraj profil fotoğrafı
Suraj17 gün önce

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

ThisMightWork profil fotoğrafı
ThisMightWork18 gün önce

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

SYNTHLEX profil fotoğrafı
SYNTHLEX18 gün önce

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

BoyardE profil fotoğrafı
BoyardE18 gün önce

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

Ibesh profil fotoğrafı
Ibesh18 gün önce

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

路克0xLUKE777crypt profil fotoğrafı
路克0xLUKE777crypt18 gün önce

@huggingface huge useful

Manny Kalavera profil fotoğrafı
Manny Kalavera18 gün önce

@huggingface Geil. Das ist sau stark.

Benzer Videolar