Загрузка видео...

Не удалось загрузить видео

На главную

When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on Hugging Face 🤗

1,063,826 просмотров • 18 дней назад •via X (Twitter)

Комментарии: 33

Фото профиля NVIDIA AI
NVIDIA AI18 дней назад

Nemotron 3 Diarization ranked #1 of 12 systems in @voicearena_ai's initial Diarization-Bench results. Its 14.72% error rate was ~24% lower than the runner-up. Here’s a four-speaker comparison from the benchmark. Full results:

Фото профиля NVIDIA AI
NVIDIA AI18 дней назад

We also put together a live demo on @huggingface if you want to try it yourself: And, a blog with more on how it works and how to get started:

Фото профиля minime · fireply.ai
minime · fireply.ai18 дней назад

@huggingface love that this went straight to huggingface instead of sitting in a paper for six months

Фото профиля Jon Taylor
Jon Taylor18 дней назад

Here is @Rekka and I giving Nemo3 Diarization a spin - Pipecat Battleships!‍ ⚓️🦜 Has been a while since I built anything with diarization and this model is really impressive. Also, super easy to slot into a Pipecat pipeline. ➡️ Code: ➡️ Full 'how it works' video:

Фото профиля Stephen Turner 🇬🇧🇺🇦
Stephen Turner 🇬🇧🇺🇦18 дней назад

@huggingface Can this be made to work on Android? If so, how? I have a Pixel 8 Pro.

Фото профиля Rompel
Rompel18 дней назад

@huggingface 100M params for 8-speaker overlap diarization is a tiny budget. pyannote chokes on overlap past 3-4 speakers in my experience, that's usually where DER spikes. Curious what overlap ratio their eval set actually has, real meetings run higher than most benchmark audio.

Фото профиля Symbioza2025 | ASA | CLM AI
Symbioza2025 | ASA | CLM AI18 дней назад

This is more important than cleaner transcripts. Speaker identity + timing + overlapping speech creates something increasingly valuable for AI systems: structured event history. As voice interfaces evolve into persistent agents, knowing what was said won't be enough. Systems will need reliable context about who said it, when, in what sequence, and what action followed. That's where diarization starts becoming part of the observability stack - not merely the transcription stack. Very interesting direction.

Фото профиля drefrajo
drefrajo18 дней назад

@huggingface now we're talking! would excited to finally see ai assistants talking with a group of people, should now be possible way more easily

Фото профиля Knowix
Knowix18 дней назад

@huggingface does the background noise have effect on it?

Фото профиля Valery
Valery18 дней назад

@huggingface Oh, could you train OpenAI on this please? 😅

Фото профиля Fab
Fab18 дней назад

@huggingface This is huge, diarization is such a clusterfest still -- OR up until yesterday i suppose - go Jensen 🙌

Фото профиля Marktechpost AI
Marktechpost AI17 дней назад

@huggingface

Фото профиля GloktaCore
GloktaCore18 дней назад

@huggingface This could be a game-changer for meetings, but I wonder how many will actually bother to use it instead of just talking over each other.

Фото профиля Aditya Kharbanda
Aditya Kharbanda18 дней назад

@huggingface @openwhispr @gabrielste1n

Фото профиля Seiji Satō
Seiji Satō18 дней назад

@Scobleizer @huggingface @__cski jfyi

Фото профиля Harsh Mishra
Harsh Mishra18 дней назад

@huggingface 100M params for 8-speaker overlap handling is surprisingly small. Curious how it degrades past 8 speakers, hard cutoff or graceful accuracy drop as more voices get added?

Фото профиля Alcreon
Alcreon18 дней назад

@huggingface 100M params handling overlapping speech is the actual flex here, overlap is where every diarization model before this fell apart

Фото профиля Rakesh Gohel 🇨🇦
Rakesh Gohel 🇨🇦18 дней назад

@huggingface Overlapping speech is where transcription gets messy. Speaker-aware AI that can track who said what in real time is a meaningful step toward reliable voice agents.

Фото профиля Emmy
Emmy17 дней назад

@huggingface 💚🚀💫

Фото профиля PENIEL
PENIEL18 дней назад

@huggingface i guess we solved diarization..

Фото профиля Panther
Panther18 дней назад

@huggingface eight people talking at once and the model is taking attendance

Фото профиля George
George18 дней назад

@huggingface Sweet. Love me some Diarrheaization.

Фото профиля Bailey Simrell
Bailey Simrell18 дней назад

@huggingface now this looks pretty interesting

Фото профиля Niko Storni 🇨🇭
Niko Storni 🇨🇭18 дней назад

@huggingface woah this is huge!

Фото профиля Robert Piazza
Robert Piazza18 дней назад

@huggingface This is awesome

Фото профиля Subhash Yadav
Subhash Yadav18 дней назад

@huggingface At 100M params this can sit right next to the ASR model instead of being a separate service. The demo says live, so what's the lag before a speaker label settles? For meeting agents, a label that flips 3 seconds later breaks "who owns this action item."

Фото профиля Suraj
Suraj17 дней назад

@huggingface eight overlapping speakers at 100m params is small for that. will check this out

Фото профиля ThisMightWork
ThisMightWork18 дней назад

Please put this through the 'three people say yep while a fourth volunteers' test. That's how an innocent transcript becomes a task assigned to the wrong person.

Фото профиля SYNTHLEX
SYNTHLEX18 дней назад

@huggingface the transcript was never the hard part, knowing which of the eight people actually agreed to the deadline was

Фото профиля BoyardE
BoyardE18 дней назад

@huggingface This is huge for people with hearing problems.... a key issue today is that making all the "right frequencies" louder just isn't helping anymore... this could help!!

Фото профиля Ibesh
Ibesh18 дней назад

@huggingface Speaker separation is the difference between a transcript you can skim and one you can actually use. A live demo makes that gap easy to judge.

Фото профиля 路克0xLUKE777crypt
路克0xLUKE777crypt18 дней назад

@huggingface huge useful

Фото профиля Manny Kalavera
Manny Kalavera18 дней назад

@huggingface Geil. Das ist sau stark.

Похожие видео