Загрузка видео...

Не удалось загрузить видео

На главную

NVIDIA just released a new open source transcription model, Nemotron Speech ASR, designed from the ground up for low-latency use cases like voice agents. Here's a voice agent built with this new model. 24ms transcription finalization and total voice-to-voice inference time under 500ms. This agent actually uses *three* NVIDIA...

274,867 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 59

Фото профиля kwindla
kwindla9 месяцев назад

Here's a technical write-up about the voice agent in the video above, the three NVIDIA models, how to deploy to production, and some fun optimizations if you're running locally on a single GPU:

Фото профиля kwindla
kwindla9 месяцев назад

Code is all here: You can deploy these models to @modal cloud really easily. (I love the Modal developer experience.) To run locally, you'll need to build a Docker container (because, you know, bleeding edge vLLM, llama.cpp, CUDA for Blackwell, etc). But the Dockerfile in the repo should "just work" on DGX Spark and RTX 5090. If you have trouble, or make patches to extend to other platforms, please let me know!

Фото профиля Rousseau & Reese Think Tank
Rousseau & Reese Think Tank9 месяцев назад

Finally I can fully automate my elderly scamming business 😌

Фото профиля Hedayat
Hedayat9 месяцев назад

The headline isn’t ASR quality. It’s that the entire voice stack is now open — ASR → LLM → TTS — with weights and training code. That’s how ecosystems form.

Фото профиля kwindla
kwindla9 месяцев назад

100%

Фото профиля Josh Garrett
Josh Garrett9 месяцев назад

Looks perfect for our new platform on Unity 6, thank you

Фото профиля Anthem
Anthem9 месяцев назад

Nano - means 30Gb for them? 😭

Фото профиля kwindla
kwindla9 месяцев назад

I mean, I actually agree with calling a model like this “nano” for what we do in the voice AI space. 30B parameters is pretty much the bare minimum size that can do open ended, human-like multi turn conversation and reliable tool calling! But I know what you mean.

Фото профиля Sadi Moodi
Sadi Moodi9 месяцев назад

this is very good, can i run the docker containers on my windows machine with RTX 4090?

Фото профиля kwindla
kwindla9 месяцев назад

Windows shouldn't be a problem, but the memory on the 4090 is not quite enough to run all three models at the same time. Definitely worth testing, though, and maybe experimenting with a smaller Nemotron 3 Nano quant. (One of the smaller 4-bit, or even a 3-bit.) If you try the container build (Dockerfile.unified) and run into problems, post an issue on the GitHub repo.

Фото профиля Sadi Moodi
Sadi Moodi9 месяцев назад

Codex is telling me i cant run nateviely on windows, only via WSL, I asked it to convert the code, lets see

Фото профиля kwindla
kwindla9 месяцев назад

That's true for Pipecat (you need WSL). But the actual Docker container should work on Windows, I would have thought.

Фото профиля Tom Osman 🐦‍⬛
Tom Osman 🐦‍⬛9 месяцев назад

fast booiiii

Фото профиля Ettore Di Giacinto
Ettore Di Giacinto9 месяцев назад

That's gonna be a neat addition for @LocalAI_API !

Фото профиля Max K
Max K9 месяцев назад

this is cool! i'd love to hear the turn detection responsiveness with this pipeline, i wonder how much more natural it'll feel

Фото профиля kwindla
kwindla9 месяцев назад

You can hear it in the demo! That’s the Pipecat open source, native audio smart turn model.

Фото профиля Damien C. Tanner
Damien C. Tanner9 месяцев назад

This looks terrific. We may deploy the sst model for @uselayercode users.

Фото профиля J.R Vale
J.R Vale9 месяцев назад

Can't wait to spin this up locally on my setup. Thanks for sharing the repo and blog!

Фото профиля Joshua March
Joshua March9 месяцев назад

Wowzer. Excited to try these out!

Фото профиля Anayat
Anayat9 месяцев назад

500ms voice to voice with three Nemotron models is wild, the real question is how far you can push turn taking and barge in before UX breaks. Sub 300ms feels like the magic number in production agents.

Фото профиля kwindla
kwindla9 месяцев назад

Hot take, but sub-300ms is *too* fast unless turn detection is *perfect*. I don’t even like talking to most actual people who do sub-300ms responses. 😜

Фото профиля Contextrix
Contextrix9 месяцев назад

Nvidia releasing truly open source models for every part of the voice pipeline is huge for developers. Achieving under 500ms voice-to-voice latency makes conversational AI feel almost instant.

Фото профиля Ditectrev
Ditectrev9 месяцев назад

Great news!🤗

Фото профиля Kurt
Kurt9 месяцев назад

It's impressive but still quite unnatural sounding.

Фото профиля kwindla
kwindla9 месяцев назад

The voice “naturalness” is really important. This is a prerelease checkpoint of the Magpie model. Think of it as a placeholder.

Фото профиля Amit Tripathi
Amit Tripathi9 месяцев назад

How does it compare with Whisper in WER for various langauges?

Фото профиля kwindla
kwindla9 месяцев назад

The general benchmark for the model in English shows a slightly better WER than Whisper V3 large in English: It is much, much faster, too, which matters a great deal for conversational voice use cases. I have not seen benchmarks for non-English.

Фото профиля Amit Tripathi
Amit Tripathi9 месяцев назад

Thanks.

Фото профиля YannickMrCrypto
YannickMrCrypto9 месяцев назад

It's really fast!

Фото профиля Juan Mugica
Juan Mugica9 месяцев назад

@eherrerosj

Фото профиля Ethan Alley
Ethan Alley9 месяцев назад

very impressive quality and speed for open source!! how much would you estimate this costs for all-in for 1 hour of real time convo?

Фото профиля ISO8601x
ISO8601x9 месяцев назад

<sigh>

Фото профиля Thomas Hill
Thomas Hill9 месяцев назад

Rapid ⚡️

Фото профиля Deva.me
Deva.me9 месяцев назад

Latency is the unlock. Once voice loops stay sub-second end to end, agents stop feeling like demos and start feeling native.

Фото профиля Shreyans Bhansali
Shreyans Bhansali9 месяцев назад

24ms is wild my brain still buffering at like 2000ms

Фото профиля kwindla
kwindla9 месяцев назад

Current-generation NVIDIA hardware is definitely faster than my meat brain, too.

Фото профиля Harry Columbo Suisse
Harry Columbo Suisse9 месяцев назад

@askcodi But your brain contains just around 90% of water, 10% other ingredients, need only a little bit food, drink to fuel it somehow and is so in terms of cost probably still lightyears ahead. It runs parallel permanent not only hear-think-speak task ..also manage few other tasks.

Фото профиля kwindla
kwindla9 месяцев назад

@askcodi My brain runs on artisanal hot sauce and expensive sushi. I think it's a tie.

Фото профиля Harry Columbo Suisse
Harry Columbo Suisse9 месяцев назад

@askcodi Hahaaa, perfect. A brain with style.

Фото профиля kwindla
kwindla9 месяцев назад

@askcodi Oh, yeah, oysters. I forgot oysters.

Фото профиля Walowitz - e/acc
Walowitz - e/acc9 месяцев назад

Latency is so important

Фото профиля Nikhil Deshpande
Nikhil Deshpande9 месяцев назад

This points to a broader shift: AI interfaces are becoming real-time, open, and infrastructure-aware. For enterprises, the advantage won’t come from owning models—but from integrating low-latency AI reliably across customer, operations, and support workflows.

Фото профиля kwindla
kwindla9 месяцев назад

I do think this next generation of open models is going to open up real opportunities to fine-tune models on internal enterprise data, which is going to be a big shift.

Фото профиля Freedom_Aint_Free
Freedom_Aint_Free9 месяцев назад

In the old sci-fi shows or even books, the voice of the AI used to be a appropriate metallic vocoder like voice, to this day I search for a voice assistant with that Kraftwerkian vocoder voice and none has it. It's just me who wants that ?

Фото профиля kwindla
kwindla9 месяцев назад

It’s funny you say that, because a couple of months ago I wanted that exactly, and all my friends who train voice model were just, “no.” All the voices are too good now. 😀

Фото профиля Freedom_Aint_Free
Freedom_Aint_Free9 месяцев назад

Yep. Vocoder voice is simply the voice of the future (or the voice of energy according to Kraftwerk !) and the history of the vocoder itself has to do with spectral analysis of the human voice and information

Фото профиля cowtower
cowtower9 месяцев назад

There are much better solutions in the TTS space. Chatterbox or Soprano come to mind. Chatterbox will give you MUCH more lifelike results, and Soprano will give you noticeably better results and still be faster than this library.

Фото профиля kwindla
kwindla9 месяцев назад

TTFB numbers or it didn’t happen. 😀

Фото профиля Stormlord
Stormlord9 месяцев назад

It has ai!

Фото профиля Manish Baghel
Manish Baghel8 месяцев назад

How does the cost of this setup compare to one like Deepgram + Gemini Flash + Cartesia?

Фото профиля Vann Snow
Vann Snow9 месяцев назад

👀

Фото профиля kyc-kyc
kyc-kyc9 месяцев назад

Scammers in India would be soooooo happy

Фото профиля letsbuildmore
letsbuildmore9 месяцев назад

Kwindla, which setup wud u recommend for local iphone voice agent?

Фото профиля kwindla
kwindla9 месяцев назад

Sadly, we are still a ways away from LLMs small enough to run on an iPhone that can do good, open-ended conversation. What’s your use case?

Фото профиля letsbuildmore
letsbuildmore9 месяцев назад

Medical. Privacy reason don’t want to send user data back to cloud LLM. On device is super important. You think Liquid LM on device LLM aren’t doing well?

Фото профиля kwindla
kwindla9 месяцев назад

We do lots of production deployments for healthcare and financial services with all the models running in a VPC. That approach is approved for regulated industries in every country I know of. You have to keep the data in the VPC and in the country of origin. But it’s the standard approach.

Фото профиля letsbuildmore
letsbuildmore9 месяцев назад

Could you share some docs/resources on this ?

Фото профиля Tech Odyssey
Tech Odyssey8 месяцев назад

This is impressive, especially the latency numbers. Sub-500ms voice to voice with fully open models is exactly what real voice agents have been missing. If NVIDIA can keep pushing open ASR and TTS this far, it really does start to feel like open source can compete on end to end experience, not just benchmarks....

Фото профиля David Im
David Im9 месяцев назад

thanks, will work on it

Похожие видео

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

333,276 просмотров • 1 месяц назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 4 месяцев назад

Nvidia has just announced Alpamayo 2 Super, an open 34 billion parameter reasoning vision-language-action model designed to accelerate the development of autonomous vehicles. This new model combines the NVIDIA Cosmos 3 Super reasoning model with a 2 billion parameter diffusion-based action expert model, and is post trained with reinforcement learning. The model can return multiple outputs: future trajectory plans, reasoning traces, grounded answers to questions about the scenes, and auto label generation. The model weights are now available for anyone to download on Hugging Face, and the inference code has been posted to GitHub. Distilled models can be deployed commercially without any further permission from Nvidia, and model outputs carry no license conditions. Automakers can distill down a compact version of this model that can run on the Nvidia computer in the car. Major kudos to Nvidia and Jensen Huang for advancing the state of the industry by releasing this as an open model with permissive licensing. Jensen isn't just paying lip service to the idea of open models, Nvidia is actually contributing to the ecosystem — and it's great for their business, because it helps sell more Thor computers that go in the car. Anyone can go download the model and play with it. If you do, let me know what you think. Personally I think it's so cool that we have open weights models that are this advanced, for anyone to download.

Whole Mars Catalog

45,595 просмотров • 2 месяцев назад