Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

NVIDIA just released a new open source transcription model, Nemotron Speech ASR, designed from the ground up for low-latency use cases like voice agents. Here's a voice agent built with this new model. 24ms transcription finalization and total voice-to-voice inference time under 500ms. This agent actually uses *three* NVIDIA...

274,867 Aufrufe • vor 9 Monaten •via X (Twitter)

59 Kommentare

Profilbild von kwindla
kwindlavor 9 Monaten

Here's a technical write-up about the voice agent in the video above, the three NVIDIA models, how to deploy to production, and some fun optimizations if you're running locally on a single GPU:

Profilbild von kwindla
kwindlavor 9 Monaten

Code is all here: You can deploy these models to @modal cloud really easily. (I love the Modal developer experience.) To run locally, you'll need to build a Docker container (because, you know, bleeding edge vLLM, llama.cpp, CUDA for Blackwell, etc). But the Dockerfile in the repo should "just work" on DGX Spark and RTX 5090. If you have trouble, or make patches to extend to other platforms, please let me know!

Profilbild von Rousseau & Reese Think Tank
Rousseau & Reese Think Tankvor 9 Monaten

Finally I can fully automate my elderly scamming business 😌

Profilbild von Hedayat
Hedayatvor 9 Monaten

The headline isn’t ASR quality. It’s that the entire voice stack is now open — ASR → LLM → TTS — with weights and training code. That’s how ecosystems form.

Profilbild von kwindla
kwindlavor 9 Monaten

100%

Profilbild von Josh Garrett
Josh Garrettvor 9 Monaten

Looks perfect for our new platform on Unity 6, thank you

Profilbild von Anthem
Anthemvor 9 Monaten

Nano - means 30Gb for them? 😭

Profilbild von kwindla
kwindlavor 9 Monaten

I mean, I actually agree with calling a model like this “nano” for what we do in the voice AI space. 30B parameters is pretty much the bare minimum size that can do open ended, human-like multi turn conversation and reliable tool calling! But I know what you mean.

Profilbild von Sadi Moodi
Sadi Moodivor 9 Monaten

this is very good, can i run the docker containers on my windows machine with RTX 4090?

Profilbild von kwindla
kwindlavor 9 Monaten

Windows shouldn't be a problem, but the memory on the 4090 is not quite enough to run all three models at the same time. Definitely worth testing, though, and maybe experimenting with a smaller Nemotron 3 Nano quant. (One of the smaller 4-bit, or even a 3-bit.) If you try the container build (Dockerfile.unified) and run into problems, post an issue on the GitHub repo.

Profilbild von Sadi Moodi
Sadi Moodivor 9 Monaten

Codex is telling me i cant run nateviely on windows, only via WSL, I asked it to convert the code, lets see

Profilbild von kwindla
kwindlavor 9 Monaten

That's true for Pipecat (you need WSL). But the actual Docker container should work on Windows, I would have thought.

Profilbild von Tom Osman 🐦‍⬛
Tom Osman 🐦‍⬛vor 9 Monaten

fast booiiii

Profilbild von Ettore Di Giacinto
Ettore Di Giacintovor 9 Monaten

That's gonna be a neat addition for @LocalAI_API !

Profilbild von Max K
Max Kvor 9 Monaten

this is cool! i'd love to hear the turn detection responsiveness with this pipeline, i wonder how much more natural it'll feel

Profilbild von kwindla
kwindlavor 9 Monaten

You can hear it in the demo! That’s the Pipecat open source, native audio smart turn model.

Profilbild von Damien C. Tanner
Damien C. Tannervor 9 Monaten

This looks terrific. We may deploy the sst model for @uselayercode users.

Profilbild von J.R Vale
J.R Valevor 9 Monaten

Can't wait to spin this up locally on my setup. Thanks for sharing the repo and blog!

Profilbild von Joshua March
Joshua Marchvor 9 Monaten

Wowzer. Excited to try these out!

Profilbild von Anayat
Anayatvor 9 Monaten

500ms voice to voice with three Nemotron models is wild, the real question is how far you can push turn taking and barge in before UX breaks. Sub 300ms feels like the magic number in production agents.

Profilbild von kwindla
kwindlavor 9 Monaten

Hot take, but sub-300ms is *too* fast unless turn detection is *perfect*. I don’t even like talking to most actual people who do sub-300ms responses. 😜

Profilbild von Contextrix
Contextrixvor 9 Monaten

Nvidia releasing truly open source models for every part of the voice pipeline is huge for developers. Achieving under 500ms voice-to-voice latency makes conversational AI feel almost instant.

Profilbild von Ditectrev
Ditectrevvor 9 Monaten

Great news!🤗

Profilbild von Kurt
Kurtvor 9 Monaten

It's impressive but still quite unnatural sounding.

Profilbild von kwindla
kwindlavor 9 Monaten

The voice “naturalness” is really important. This is a prerelease checkpoint of the Magpie model. Think of it as a placeholder.

Profilbild von Amit Tripathi
Amit Tripathivor 9 Monaten

How does it compare with Whisper in WER for various langauges?

Profilbild von kwindla
kwindlavor 9 Monaten

The general benchmark for the model in English shows a slightly better WER than Whisper V3 large in English: It is much, much faster, too, which matters a great deal for conversational voice use cases. I have not seen benchmarks for non-English.

Profilbild von Amit Tripathi
Amit Tripathivor 9 Monaten

Thanks.

Profilbild von YannickMrCrypto
YannickMrCryptovor 9 Monaten

It's really fast!

Profilbild von Juan Mugica
Juan Mugicavor 9 Monaten

@eherrerosj

Profilbild von Ethan Alley
Ethan Alleyvor 9 Monaten

very impressive quality and speed for open source!! how much would you estimate this costs for all-in for 1 hour of real time convo?

Profilbild von ISO8601x
ISO8601xvor 9 Monaten

<sigh>

Profilbild von Thomas Hill
Thomas Hillvor 9 Monaten

Rapid ⚡️

Profilbild von Deva.me
Deva.mevor 9 Monaten

Latency is the unlock. Once voice loops stay sub-second end to end, agents stop feeling like demos and start feeling native.

Profilbild von Shreyans Bhansali
Shreyans Bhansalivor 9 Monaten

24ms is wild my brain still buffering at like 2000ms

Profilbild von kwindla
kwindlavor 9 Monaten

Current-generation NVIDIA hardware is definitely faster than my meat brain, too.

Profilbild von Harry Columbo Suisse
Harry Columbo Suissevor 9 Monaten

@askcodi But your brain contains just around 90% of water, 10% other ingredients, need only a little bit food, drink to fuel it somehow and is so in terms of cost probably still lightyears ahead. It runs parallel permanent not only hear-think-speak task ..also manage few other tasks.

Profilbild von kwindla
kwindlavor 9 Monaten

@askcodi My brain runs on artisanal hot sauce and expensive sushi. I think it's a tie.

Profilbild von Harry Columbo Suisse
Harry Columbo Suissevor 9 Monaten

@askcodi Hahaaa, perfect. A brain with style.

Profilbild von kwindla
kwindlavor 9 Monaten

@askcodi Oh, yeah, oysters. I forgot oysters.

Profilbild von Walowitz - e/acc
Walowitz - e/accvor 9 Monaten

Latency is so important

Profilbild von Nikhil Deshpande
Nikhil Deshpandevor 9 Monaten

This points to a broader shift: AI interfaces are becoming real-time, open, and infrastructure-aware. For enterprises, the advantage won’t come from owning models—but from integrating low-latency AI reliably across customer, operations, and support workflows.

Profilbild von kwindla
kwindlavor 9 Monaten

I do think this next generation of open models is going to open up real opportunities to fine-tune models on internal enterprise data, which is going to be a big shift.

Profilbild von Freedom_Aint_Free
Freedom_Aint_Freevor 9 Monaten

In the old sci-fi shows or even books, the voice of the AI used to be a appropriate metallic vocoder like voice, to this day I search for a voice assistant with that Kraftwerkian vocoder voice and none has it. It's just me who wants that ?

Profilbild von kwindla
kwindlavor 9 Monaten

It’s funny you say that, because a couple of months ago I wanted that exactly, and all my friends who train voice model were just, “no.” All the voices are too good now. 😀

Profilbild von Freedom_Aint_Free
Freedom_Aint_Freevor 9 Monaten

Yep. Vocoder voice is simply the voice of the future (or the voice of energy according to Kraftwerk !) and the history of the vocoder itself has to do with spectral analysis of the human voice and information

Profilbild von cowtower
cowtowervor 9 Monaten

There are much better solutions in the TTS space. Chatterbox or Soprano come to mind. Chatterbox will give you MUCH more lifelike results, and Soprano will give you noticeably better results and still be faster than this library.

Profilbild von kwindla
kwindlavor 9 Monaten

TTFB numbers or it didn’t happen. 😀

Profilbild von Stormlord
Stormlordvor 9 Monaten

It has ai!

Profilbild von Manish Baghel
Manish Baghelvor 8 Monaten

How does the cost of this setup compare to one like Deepgram + Gemini Flash + Cartesia?

Profilbild von Vann Snow
Vann Snowvor 9 Monaten

👀

Profilbild von kyc-kyc
kyc-kycvor 9 Monaten

Scammers in India would be soooooo happy

Profilbild von letsbuildmore
letsbuildmorevor 9 Monaten

Kwindla, which setup wud u recommend for local iphone voice agent?

Profilbild von kwindla
kwindlavor 9 Monaten

Sadly, we are still a ways away from LLMs small enough to run on an iPhone that can do good, open-ended conversation. What’s your use case?

Profilbild von letsbuildmore
letsbuildmorevor 9 Monaten

Medical. Privacy reason don’t want to send user data back to cloud LLM. On device is super important. You think Liquid LM on device LLM aren’t doing well?

Profilbild von kwindla
kwindlavor 9 Monaten

We do lots of production deployments for healthcare and financial services with all the models running in a VPC. That approach is approved for regulated industries in every country I know of. You have to keep the data in the VPC and in the country of origin. But it’s the standard approach.

Profilbild von letsbuildmore
letsbuildmorevor 9 Monaten

Could you share some docs/resources on this ?

Profilbild von Tech Odyssey
Tech Odysseyvor 8 Monaten

This is impressive, especially the latency numbers. Sub-500ms voice to voice with fully open models is exactly what real voice agents have been missing. If NVIDIA can keep pushing open ASR and TTS this far, it really does start to feel like open source can compete on end to end experience, not just benchmarks....

Profilbild von David Im
David Imvor 9 Monaten

thanks, will work on it

Ähnliche Videos

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

333,276 Aufrufe • vor 1 Monat

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 4 Monaten