Video wird geladen...
Video konnte nicht geladen werden
Listen, speak, handle interruptions, and call tools in one live conversation. ๐๏ธ Introducing NVIDIA-NemotronLabs-VoiceChat-11B, NVIDIAโs end-to-end full-duplex model for real-time voice agents. ๐ค ๐ It ranks #2 among open full-duplex models on both VoiceBench and Full-Duplex-Bench 1.0. โก Natural turn-taking responds in ~448 ms, while barge-in lets users interrupt... show more
22,424 Aufrufe โข vor 1 Monat โขvia X (Twitter)
7 Kommentare

Real time voice that actually listens. 448ms responses, live tool calls, zero awkward pauses.

Everything good but the license bites ๐

Wow, that's super impressive! Sounds like a game-changer for voice agents. Congrats on the ranking!

it has pretty high latency for a duplex model for some reason. it doesn't sound like 500ms response latency, more like >1.5-2s latency. awkward...

Full-duplex fixes the speed. The catch is the glass box: a cascade gives you a clean per-utterance record you can audit, duplex blurs who said what and when. Getting duplex latency without losing that structure is the actual hard part.

NVIDIA's new end-to-end voice model brings seamless conversations and live tool use together.

Barge-in latency that low finally makes voice agents feel real.
