Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

NVIDIA just released a new open source transcription model, Nemotron Speech ASR, designed from the ground up for low-latency use cases like voice agents. Here's a voice agent built with this new model. 24ms transcription finalization and total voice-to-voice inference time under 500ms. This agent actually uses *three* NVIDIA...

274,867 görüntüleme • 9 ay önce •via X (Twitter)

59 Yorum

kwindla profil fotoğrafı
kwindla9 ay önce

Here's a technical write-up about the voice agent in the video above, the three NVIDIA models, how to deploy to production, and some fun optimizations if you're running locally on a single GPU:

kwindla profil fotoğrafı
kwindla9 ay önce

Code is all here: You can deploy these models to @modal cloud really easily. (I love the Modal developer experience.) To run locally, you'll need to build a Docker container (because, you know, bleeding edge vLLM, llama.cpp, CUDA for Blackwell, etc). But the Dockerfile in the repo should "just work" on DGX Spark and RTX 5090. If you have trouble, or make patches to extend to other platforms, please let me know!

Rousseau & Reese Think Tank profil fotoğrafı
Rousseau & Reese Think Tank9 ay önce

Finally I can fully automate my elderly scamming business 😌

Hedayat profil fotoğrafı
Hedayat9 ay önce

The headline isn’t ASR quality. It’s that the entire voice stack is now open — ASR → LLM → TTS — with weights and training code. That’s how ecosystems form.

kwindla profil fotoğrafı
kwindla9 ay önce

100%

Josh Garrett profil fotoğrafı
Josh Garrett9 ay önce

Looks perfect for our new platform on Unity 6, thank you

Anthem profil fotoğrafı
Anthem9 ay önce

Nano - means 30Gb for them? 😭

kwindla profil fotoğrafı
kwindla9 ay önce

I mean, I actually agree with calling a model like this “nano” for what we do in the voice AI space. 30B parameters is pretty much the bare minimum size that can do open ended, human-like multi turn conversation and reliable tool calling! But I know what you mean.

Sadi Moodi profil fotoğrafı
Sadi Moodi9 ay önce

this is very good, can i run the docker containers on my windows machine with RTX 4090?

kwindla profil fotoğrafı
kwindla9 ay önce

Windows shouldn't be a problem, but the memory on the 4090 is not quite enough to run all three models at the same time. Definitely worth testing, though, and maybe experimenting with a smaller Nemotron 3 Nano quant. (One of the smaller 4-bit, or even a 3-bit.) If you try the container build (Dockerfile.unified) and run into problems, post an issue on the GitHub repo.

Sadi Moodi profil fotoğrafı
Sadi Moodi9 ay önce

Codex is telling me i cant run nateviely on windows, only via WSL, I asked it to convert the code, lets see

kwindla profil fotoğrafı
kwindla9 ay önce

That's true for Pipecat (you need WSL). But the actual Docker container should work on Windows, I would have thought.

Tom Osman 🐦‍⬛ profil fotoğrafı
Tom Osman 🐦‍⬛9 ay önce

fast booiiii

Ettore Di Giacinto profil fotoğrafı
Ettore Di Giacinto9 ay önce

That's gonna be a neat addition for @LocalAI_API !

Max K profil fotoğrafı
Max K9 ay önce

this is cool! i'd love to hear the turn detection responsiveness with this pipeline, i wonder how much more natural it'll feel

kwindla profil fotoğrafı
kwindla9 ay önce

You can hear it in the demo! That’s the Pipecat open source, native audio smart turn model.

Damien C. Tanner profil fotoğrafı
Damien C. Tanner9 ay önce

This looks terrific. We may deploy the sst model for @uselayercode users.

J.R Vale profil fotoğrafı
J.R Vale9 ay önce

Can't wait to spin this up locally on my setup. Thanks for sharing the repo and blog!

Joshua March profil fotoğrafı
Joshua March9 ay önce

Wowzer. Excited to try these out!

Anayat profil fotoğrafı
Anayat9 ay önce

500ms voice to voice with three Nemotron models is wild, the real question is how far you can push turn taking and barge in before UX breaks. Sub 300ms feels like the magic number in production agents.

kwindla profil fotoğrafı
kwindla9 ay önce

Hot take, but sub-300ms is *too* fast unless turn detection is *perfect*. I don’t even like talking to most actual people who do sub-300ms responses. 😜

Contextrix profil fotoğrafı
Contextrix9 ay önce

Nvidia releasing truly open source models for every part of the voice pipeline is huge for developers. Achieving under 500ms voice-to-voice latency makes conversational AI feel almost instant.

Ditectrev profil fotoğrafı
Ditectrev9 ay önce

Great news!🤗

Kurt profil fotoğrafı
Kurt9 ay önce

It's impressive but still quite unnatural sounding.

kwindla profil fotoğrafı
kwindla9 ay önce

The voice “naturalness” is really important. This is a prerelease checkpoint of the Magpie model. Think of it as a placeholder.

Amit Tripathi profil fotoğrafı
Amit Tripathi9 ay önce

How does it compare with Whisper in WER for various langauges?

kwindla profil fotoğrafı
kwindla9 ay önce

The general benchmark for the model in English shows a slightly better WER than Whisper V3 large in English: It is much, much faster, too, which matters a great deal for conversational voice use cases. I have not seen benchmarks for non-English.

Amit Tripathi profil fotoğrafı
Amit Tripathi9 ay önce

Thanks.

YannickMrCrypto profil fotoğrafı
YannickMrCrypto9 ay önce

It's really fast!

Juan Mugica profil fotoğrafı
Juan Mugica9 ay önce

@eherrerosj

Ethan Alley profil fotoğrafı
Ethan Alley9 ay önce

very impressive quality and speed for open source!! how much would you estimate this costs for all-in for 1 hour of real time convo?

ISO8601x profil fotoğrafı
ISO8601x9 ay önce

<sigh>

Thomas Hill profil fotoğrafı
Thomas Hill9 ay önce

Rapid ⚡️

Deva.me profil fotoğrafı
Deva.me9 ay önce

Latency is the unlock. Once voice loops stay sub-second end to end, agents stop feeling like demos and start feeling native.

Shreyans Bhansali profil fotoğrafı
Shreyans Bhansali9 ay önce

24ms is wild my brain still buffering at like 2000ms

kwindla profil fotoğrafı
kwindla9 ay önce

Current-generation NVIDIA hardware is definitely faster than my meat brain, too.

Harry Columbo Suisse profil fotoğrafı
Harry Columbo Suisse9 ay önce

@askcodi But your brain contains just around 90% of water, 10% other ingredients, need only a little bit food, drink to fuel it somehow and is so in terms of cost probably still lightyears ahead. It runs parallel permanent not only hear-think-speak task ..also manage few other tasks.

kwindla profil fotoğrafı
kwindla9 ay önce

@askcodi My brain runs on artisanal hot sauce and expensive sushi. I think it's a tie.

Harry Columbo Suisse profil fotoğrafı
Harry Columbo Suisse9 ay önce

@askcodi Hahaaa, perfect. A brain with style.

kwindla profil fotoğrafı
kwindla9 ay önce

@askcodi Oh, yeah, oysters. I forgot oysters.

Walowitz - e/acc profil fotoğrafı
Walowitz - e/acc9 ay önce

Latency is so important

Nikhil Deshpande profil fotoğrafı
Nikhil Deshpande9 ay önce

This points to a broader shift: AI interfaces are becoming real-time, open, and infrastructure-aware. For enterprises, the advantage won’t come from owning models—but from integrating low-latency AI reliably across customer, operations, and support workflows.

kwindla profil fotoğrafı
kwindla9 ay önce

I do think this next generation of open models is going to open up real opportunities to fine-tune models on internal enterprise data, which is going to be a big shift.

Freedom_Aint_Free profil fotoğrafı
Freedom_Aint_Free9 ay önce

In the old sci-fi shows or even books, the voice of the AI used to be a appropriate metallic vocoder like voice, to this day I search for a voice assistant with that Kraftwerkian vocoder voice and none has it. It's just me who wants that ?

kwindla profil fotoğrafı
kwindla9 ay önce

It’s funny you say that, because a couple of months ago I wanted that exactly, and all my friends who train voice model were just, “no.” All the voices are too good now. 😀

Freedom_Aint_Free profil fotoğrafı
Freedom_Aint_Free9 ay önce

Yep. Vocoder voice is simply the voice of the future (or the voice of energy according to Kraftwerk !) and the history of the vocoder itself has to do with spectral analysis of the human voice and information

cowtower profil fotoğrafı
cowtower9 ay önce

There are much better solutions in the TTS space. Chatterbox or Soprano come to mind. Chatterbox will give you MUCH more lifelike results, and Soprano will give you noticeably better results and still be faster than this library.

kwindla profil fotoğrafı
kwindla9 ay önce

TTFB numbers or it didn’t happen. 😀

Stormlord profil fotoğrafı
Stormlord9 ay önce

It has ai!

Manish Baghel profil fotoğrafı
Manish Baghel8 ay önce

How does the cost of this setup compare to one like Deepgram + Gemini Flash + Cartesia?

Vann Snow profil fotoğrafı
Vann Snow9 ay önce

👀

kyc-kyc profil fotoğrafı
kyc-kyc9 ay önce

Scammers in India would be soooooo happy

letsbuildmore profil fotoğrafı
letsbuildmore9 ay önce

Kwindla, which setup wud u recommend for local iphone voice agent?

kwindla profil fotoğrafı
kwindla9 ay önce

Sadly, we are still a ways away from LLMs small enough to run on an iPhone that can do good, open-ended conversation. What’s your use case?

letsbuildmore profil fotoğrafı
letsbuildmore9 ay önce

Medical. Privacy reason don’t want to send user data back to cloud LLM. On device is super important. You think Liquid LM on device LLM aren’t doing well?

kwindla profil fotoğrafı
kwindla9 ay önce

We do lots of production deployments for healthcare and financial services with all the models running in a VPC. That approach is approved for regulated industries in every country I know of. You have to keep the data in the VPC and in the country of origin. But it’s the standard approach.

letsbuildmore profil fotoğrafı
letsbuildmore9 ay önce

Could you share some docs/resources on this ?

Tech Odyssey profil fotoğrafı
Tech Odyssey8 ay önce

This is impressive, especially the latency numbers. Sub-500ms voice to voice with fully open models is exactly what real voice agents have been missing. If NVIDIA can keep pushing open ASR and TTS this far, it really does start to feel like open source can compete on end to end experience, not just benchmarks....

David Im profil fotoğrafı
David Im9 ay önce

thanks, will work on it

Benzer Videolar

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

333,276 görüntüleme • 1 ay önce

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 görüntüleme • 4 ay önce