Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Listen, speak, handle interruptions, and call tools in one live conversation. ๐ŸŽ™๏ธ Introducing NVIDIA-NemotronLabs-VoiceChat-11B, NVIDIAโ€™s end-to-end full-duplex model for real-time voice agents. ๐Ÿค– ๐Ÿ† It ranks #2 among open full-duplex models on both VoiceBench and Full-Duplex-Bench 1.0. โšก Natural turn-taking responds in ~448 ms, while barge-in lets users interrupt...

22,424 Aufrufe โ€ข vor 1 Monat โ€ขvia X (Twitter)

7 Kommentare

Profilbild von Alice The Ai Expert
Alice The Ai Expertvor 1 Monat

Real time voice that actually listens. 448ms responses, live tool calls, zero awkward pauses.

Profilbild von Mahimai Raja J โ€Ž
Mahimai Raja J โ€Žvor 1 Monat

Everything good but the license bites ๐Ÿ˜ƒ

Profilbild von Polash Khan
Polash Khanvor 1 Monat

Wow, that's super impressive! Sounds like a game-changer for voice agents. Congrats on the ranking!

Profilbild von Akmal Sultanov
Akmal Sultanovvor 1 Monat

it has pretty high latency for a duplex model for some reason. it doesn't sound like 500ms response latency, more like >1.5-2s latency. awkward...

Profilbild von Karan Singla
Karan Singlavor 1 Monat

Full-duplex fixes the speed. The catch is the glass box: a cascade gives you a clean per-utterance record you can audit, duplex blurs who said what and when. Getting duplex latency without losing that structure is the actual hard part.

Profilbild von RAZA | AI EXPLORER
RAZA | AI EXPLORERvor 1 Monat

NVIDIA's new end-to-end voice model brings seamless conversations and live tool use together.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Barge-in latency that low finally makes voice agents feel real.

ร„hnliche Videos

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

331,175 Aufrufe โ€ข vor 28 Tagen

What if your voice AI could interrupt you the moment it figured out your question - sometimes even before you finished asking it? Last week, I sat down with Neil, CEO of Gradium and co-founder of Kyutai , to talk about the future of speech-to-speech models and why he believes today's cascaded voice systems will soon look "archaic and brittle." Some highlights from our conversation: ๐ŸŽฏ How Kyutai built Moshiโ€”a full duplex conversational AI with "negative latency"โ€”in 6 months with just 4-6 people (while big tech teams had 10-20x the resources) ๐Ÿง  Why speech-to-speech models lose intelligence compared to their text counterparts (and what's being done about it) ๐Ÿ“ฑ Pocket TTS: The first voice cloning model that runs on your phone's CPUโ€”not GPU, CPU ๐Ÿค– Why robotics and spatial audio represent the next frontier (hint: current voice systems completely break in these environments) ๐Ÿ‘ถ The efficiency gap: Babies learn to speak fluently from <5,000 hours of audio. Current models train on millions of hours. We're doing something wrong. My favorite vision from Neil? The first truly contrarian AI that interrupts you mid-sentence to tell you why you're wrong. Not just more natural conversationโ€”but actually useful for testing ideas and playing devil's advocate. Full episode and detailed blog post linked in the comments ๐Ÿ‘‡ What's your take - will speech-to-speech replace cascaded systems, or will modularity keep cascaded architectures dominant even as naturalness improves?

Brooke Hopkins

13,000 Aufrufe โ€ข vor 7 Monaten

You can now generate real-time speech that sounds conversational. Microsoft just open-sourced VibeVoice, a real-time text-to-speech system with ~300 ms first audio latency and streaming input. It handles long conversations without falling apart. ๐—ง๐—ต๐—ถ๐˜€ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐˜๐—ฒ๐˜€ ๐—น๐—ผ๐—ป๐—ด, ๐—บ๐˜‚๐—น๐˜๐—ถ-๐˜€๐—ฝ๐—ฒ๐—ฎ๐—ธ๐—ฒ๐—ฟ ๐˜€๐—ฝ๐—ฒ๐—ฒ๐—ฐ๐—ต. It produces up to 90 minutes of audio. It supports up to four distinct speakers. Turn-taking stays consistent over long sessions. ๐—œ๐˜ ๐˜„๐—ผ๐—ฟ๐—ธ๐˜€ ๐—ฏ๐˜† ๐—ฟ๐—ฒ๐—ฑ๐˜‚๐—ฐ๐—ถ๐—ป๐—ด ๐˜๐—ถ๐—บ๐—ฒ ๐—ฟ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป. Audio compresses into semantic and acoustic tokens. They run at 7.5 Hz instead of frame-level audio. A language model predicts structure. A diffusion head restores acoustic detail. ๐—œ๐˜ ๐—ฎ๐—น๐—น๐—ผ๐˜„๐˜€ ๐—น๐—ผ๐˜„-๐—น๐—ฎ๐˜๐—ฒ๐—ป๐—ฐ๐˜† ๐˜€๐˜๐—ฟ๐—ฒ๐—ฎ๐—บ๐—ถ๐—ป๐—ด ๐—ฎ๐˜‚๐—ฑ๐—ถ๐—ผ. The real-time variant streams text incrementally. First speech arrives in ~300 ms. A WebSocket demo shows live generation. The code is MIT-licensed and research-only. The repo already passed 20k GitHub stars.

Lior Alexander

61,122 Aufrufe โ€ข vor 8 Monaten

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe โ€ข vor 4 Monaten