Video wird geladen...
Video konnte nicht geladen werden
NVIDIA just released a new open source transcription model, Nemotron Speech ASR, designed from the ground up for low-latency use cases like voice agents. Here's a voice agent built with this new model. 24ms transcription finalization and total voice-to-voice inference time under 500ms. This agent actually uses *three* NVIDIA... show more
274,867 Aufrufe • vor 9 Monaten •via X (Twitter)
59 Kommentare

Here's a technical write-up about the voice agent in the video above, the three NVIDIA models, how to deploy to production, and some fun optimizations if you're running locally on a single GPU:

Code is all here: You can deploy these models to @modal cloud really easily. (I love the Modal developer experience.) To run locally, you'll need to build a Docker container (because, you know, bleeding edge vLLM, llama.cpp, CUDA for Blackwell, etc). But the Dockerfile in the repo should "just work" on DGX Spark and RTX 5090. If you have trouble, or make patches to extend to other platforms, please let me know!

Finally I can fully automate my elderly scamming business 😌

The headline isn’t ASR quality. It’s that the entire voice stack is now open — ASR → LLM → TTS — with weights and training code. That’s how ecosystems form.

100%

Looks perfect for our new platform on Unity 6, thank you

Nano - means 30Gb for them? 😭

I mean, I actually agree with calling a model like this “nano” for what we do in the voice AI space. 30B parameters is pretty much the bare minimum size that can do open ended, human-like multi turn conversation and reliable tool calling! But I know what you mean.

this is very good, can i run the docker containers on my windows machine with RTX 4090?

Windows shouldn't be a problem, but the memory on the 4090 is not quite enough to run all three models at the same time. Definitely worth testing, though, and maybe experimenting with a smaller Nemotron 3 Nano quant. (One of the smaller 4-bit, or even a 3-bit.) If you try the container build (Dockerfile.unified) and run into problems, post an issue on the GitHub repo.

Codex is telling me i cant run nateviely on windows, only via WSL, I asked it to convert the code, lets see

That's true for Pipecat (you need WSL). But the actual Docker container should work on Windows, I would have thought.

fast booiiii

That's gonna be a neat addition for @LocalAI_API !

this is cool! i'd love to hear the turn detection responsiveness with this pipeline, i wonder how much more natural it'll feel

You can hear it in the demo! That’s the Pipecat open source, native audio smart turn model.

This looks terrific. We may deploy the sst model for @uselayercode users.

Can't wait to spin this up locally on my setup. Thanks for sharing the repo and blog!

Wowzer. Excited to try these out!

500ms voice to voice with three Nemotron models is wild, the real question is how far you can push turn taking and barge in before UX breaks. Sub 300ms feels like the magic number in production agents.

Hot take, but sub-300ms is *too* fast unless turn detection is *perfect*. I don’t even like talking to most actual people who do sub-300ms responses. 😜

Nvidia releasing truly open source models for every part of the voice pipeline is huge for developers. Achieving under 500ms voice-to-voice latency makes conversational AI feel almost instant.

Great news!🤗

It's impressive but still quite unnatural sounding.

The voice “naturalness” is really important. This is a prerelease checkpoint of the Magpie model. Think of it as a placeholder.

How does it compare with Whisper in WER for various langauges?

The general benchmark for the model in English shows a slightly better WER than Whisper V3 large in English: It is much, much faster, too, which matters a great deal for conversational voice use cases. I have not seen benchmarks for non-English.

Thanks.

It's really fast!

@eherrerosj

very impressive quality and speed for open source!! how much would you estimate this costs for all-in for 1 hour of real time convo?

<sigh>

Rapid ⚡️

Latency is the unlock. Once voice loops stay sub-second end to end, agents stop feeling like demos and start feeling native.

24ms is wild my brain still buffering at like 2000ms

Current-generation NVIDIA hardware is definitely faster than my meat brain, too.

@askcodi But your brain contains just around 90% of water, 10% other ingredients, need only a little bit food, drink to fuel it somehow and is so in terms of cost probably still lightyears ahead. It runs parallel permanent not only hear-think-speak task ..also manage few other tasks.

@askcodi My brain runs on artisanal hot sauce and expensive sushi. I think it's a tie.

@askcodi Hahaaa, perfect. A brain with style.

@askcodi Oh, yeah, oysters. I forgot oysters.

Latency is so important

This points to a broader shift: AI interfaces are becoming real-time, open, and infrastructure-aware. For enterprises, the advantage won’t come from owning models—but from integrating low-latency AI reliably across customer, operations, and support workflows.

I do think this next generation of open models is going to open up real opportunities to fine-tune models on internal enterprise data, which is going to be a big shift.

In the old sci-fi shows or even books, the voice of the AI used to be a appropriate metallic vocoder like voice, to this day I search for a voice assistant with that Kraftwerkian vocoder voice and none has it. It's just me who wants that ?

It’s funny you say that, because a couple of months ago I wanted that exactly, and all my friends who train voice model were just, “no.” All the voices are too good now. 😀

Yep. Vocoder voice is simply the voice of the future (or the voice of energy according to Kraftwerk !) and the history of the vocoder itself has to do with spectral analysis of the human voice and information

There are much better solutions in the TTS space. Chatterbox or Soprano come to mind. Chatterbox will give you MUCH more lifelike results, and Soprano will give you noticeably better results and still be faster than this library.

TTFB numbers or it didn’t happen. 😀

It has ai!

How does the cost of this setup compare to one like Deepgram + Gemini Flash + Cartesia?

👀

Scammers in India would be soooooo happy

Kwindla, which setup wud u recommend for local iphone voice agent?

Sadly, we are still a ways away from LLMs small enough to run on an iPhone that can do good, open-ended conversation. What’s your use case?

Medical. Privacy reason don’t want to send user data back to cloud LLM. On device is super important. You think Liquid LM on device LLM aren’t doing well?

We do lots of production deployments for healthcare and financial services with all the models running in a VPC. That approach is approved for regulated industries in every country I know of. You have to keep the data in the VPC and in the country of origin. But it’s the standard approach.

Could you share some docs/resources on this ?

This is impressive, especially the latency numbers. Sub-500ms voice to voice with fully open models is exactly what real voice agents have been missing. If NVIDIA can keep pushing open ASR and TTS this far, it really does start to feel like open source can compete on end to end experience, not just benchmarks....

thanks, will work on it

