Loading video...

Video Failed to Load

Go Home

Voice AI without the wait! ⏱️ Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for voice AI at ultra-fast inference speeds. Add it to a fully open-source, cascaded speech-to-speech stack that can be used to power existing voice apps! 🗣️

161,360 views • 1 month ago •via X (Twitter)

33 Comments

Google Gemma's profile picture
Google Gemma1 month ago

Try the demo by @andimarafioti and team: Read the blog:

Arsh's profile picture
Arsh1 month ago

tried it in an all local setup on 8GB Vram. still worked great! (had to downgrade to a 4B model, and kokoro tts though)

RGBSays's profile picture
RGBSays1 month ago

I did the same thing with the Gemma 4 12B fp4, Kokoro, and OpenWebUI 😁

Francesco Pistillo's profile picture
Francesco Pistillo1 month ago

@ollama any plan to add this to your cloud? :)

Inflectiv AI ⧉'s profile picture
Inflectiv AI ⧉1 month ago

🔥🔥

Joe Xu's profile picture
Joe Xu1 month ago

did i just watch a love, death, +robots episode?

Ritesh Chavan's profile picture
Ritesh Chavan1 month ago

Impressive 🤩

sinain's profile picture
sinain1 month ago

One of these apps is Sinain AR helper - . The AI you can video call to request any assistance.

Anthony's profile picture
Anthony1 month ago

I was about to hate and say it's not that fast. But it's actually pretty fast.

Shashank's profile picture
Shashank1 month ago

@PrakharOjha4

Vanar's profile picture
Vanar1 month ago

🔥🔥

🇸​🇭​🇪​🇷​🇱​🇴​🇨​🇰​'s profile picture
🇸​🇭​🇪​🇷​🇱​🇴​🇨​🇰​1 month ago

@openwhispr

Yujin Huang's profile picture
Yujin Huang1 month ago

@_akhaliq Amazing breakthrough! What about the security side? Check it out here.

DanieR's profile picture
DanieR1 month ago

Impressive latency breakthrough! 🚀 Gemma 4 + Cerebras as the speech-to-speech brain is a smart stack for real-time voice AI applications. Looking forward to testing the open-source pipeline. 🔊

AI and Robots's profile picture
AI and Robots1 month ago

Like this alot. Great work guys

Mert · AI Architect's profile picture
Mert · AI Architect1 month ago

cascaded s2s on Cerebras gets the latency floor under 1s. that's the threshold where voice agents stop feeling like a phone tree.

Ayush Bhattacharya's profile picture
Ayush Bhattacharya1 month ago

we run Gemma 4 31B on Cerebras in production. the speed is real, but in a cascaded stack the latency shifts to accurate end of turn detection to avoid over eagerness once the model is this fast.

Veer Kheterpal's profile picture
Veer Kheterpal1 month ago

At 1,500 tok/s, a 150-token answer is only 100 ms of decode. The model is no longer the latency bottleneck. Speech recognition, first-token latency, networking, and synthesis now dominate. Inference acceleration does more than improve the experience. It exposes the rest of the stack.

Igor Ilyinsky JoinListenUp.com rogi.eth's profile picture
Igor Ilyinsky JoinListenUp.com rogi.eth1 month ago

Voice isn’t an app. It’s the interface.

Mike Gannotti's profile picture
Mike Gannotti1 month ago

Nice!!

Leon - construindo apps's profile picture
Leon - construindo apps1 month ago

@grok Consigo rodar isso no celular?

Neel Samadder's profile picture
Neel Samadder1 month ago

Fast voice changes the UX more than people expect. Once the pause feels human, users stop treating the agent like a form and start treating it like a teammate.

Kiran Adimatyam's profile picture
Kiran Adimatyam1 month ago

This is great work. Very impressive on how fast this is.

Armando Cesar's profile picture
Armando Cesar1 month ago

latency is the whole game for voice, under ~300ms round trip it stops feeling like a walkie talkie. curious how it holds up with tool calls in the loop tho, thats where every voice stack ive tried falls apart

Lyes Sbahi's profile picture
Lyes Sbahi1 month ago

@huggingface Je l’ai testé et effectivement c’est impressionnant en local

Levieu's profile picture
Levieu1 month ago

Ok

Carlos Eduado Gómez (Tato)'s profile picture
Carlos Eduado Gómez (Tato)1 month ago

@huggingface Awesome! Let’s test it out..

Aiden Milan's profile picture
Aiden Milan1 month ago

Leveraging Gemma 4 31B on Cerebras infrastructure for ultra-fast, open-source cascaded speech-to-speech AI is a game-changer for voice applications! Near-zero latency is critical for conversational AI to feel natural. What is the average end-to-end latency achieved with this stack?

boring ai things's profile picture
boring ai things1 month ago

there is literally 1-2s delay XD

#INTERNETofAGENTS's profile picture
#INTERNETofAGENTS1 month ago

Digital naps. 🫡

Noju's profile picture
Noju1 month ago

Crazy speed! But Not all Tools replies fast. Does the task blocks the Model? Eg, if the task runs for 1minute, can i still talk to the model?. Or Is this Application Layer thing?

Sri's profile picture
Sri1 month ago

I believe it passes this test as it’s a reasoning model

Mena Botrous's profile picture
Mena Botrous1 month ago

Fast open inference makes natural voice UX viable without vendor lock-in.

Related Videos