Loading video...
Video Failed to Load
Voice AI without the wait! ⏱️ Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for voice AI at ultra-fast inference speeds. Add it to a fully open-source, cascaded speech-to-speech stack that can be used to power existing voice apps! 🗣️
161,360 views • 1 month ago •via X (Twitter)
33 Comments

Try the demo by @andimarafioti and team: Read the blog:

tried it in an all local setup on 8GB Vram. still worked great! (had to downgrade to a 4B model, and kokoro tts though)

I did the same thing with the Gemma 4 12B fp4, Kokoro, and OpenWebUI 😁

@ollama any plan to add this to your cloud? :)

🔥🔥

did i just watch a love, death, +robots episode?

Impressive 🤩

One of these apps is Sinain AR helper - . The AI you can video call to request any assistance.

I was about to hate and say it's not that fast. But it's actually pretty fast.

@PrakharOjha4

🔥🔥

@openwhispr

@_akhaliq Amazing breakthrough! What about the security side? Check it out here.

Impressive latency breakthrough! 🚀 Gemma 4 + Cerebras as the speech-to-speech brain is a smart stack for real-time voice AI applications. Looking forward to testing the open-source pipeline. 🔊

Like this alot. Great work guys

cascaded s2s on Cerebras gets the latency floor under 1s. that's the threshold where voice agents stop feeling like a phone tree.

we run Gemma 4 31B on Cerebras in production. the speed is real, but in a cascaded stack the latency shifts to accurate end of turn detection to avoid over eagerness once the model is this fast.

At 1,500 tok/s, a 150-token answer is only 100 ms of decode. The model is no longer the latency bottleneck. Speech recognition, first-token latency, networking, and synthesis now dominate. Inference acceleration does more than improve the experience. It exposes the rest of the stack.

Voice isn’t an app. It’s the interface.

Nice!!

@grok Consigo rodar isso no celular?

Fast voice changes the UX more than people expect. Once the pause feels human, users stop treating the agent like a form and start treating it like a teammate.

This is great work. Very impressive on how fast this is.

latency is the whole game for voice, under ~300ms round trip it stops feeling like a walkie talkie. curious how it holds up with tool calls in the loop tho, thats where every voice stack ive tried falls apart

@huggingface Je l’ai testé et effectivement c’est impressionnant en local

Ok

@huggingface Awesome! Let’s test it out..

Leveraging Gemma 4 31B on Cerebras infrastructure for ultra-fast, open-source cascaded speech-to-speech AI is a game-changer for voice applications! Near-zero latency is critical for conversational AI to feel natural. What is the average end-to-end latency achieved with this stack?

there is literally 1-2s delay XD

Digital naps. 🫡

Crazy speed! But Not all Tools replies fast. Does the task blocks the Model? Eg, if the task runs for 1minute, can i still talk to the model?. Or Is this Application Layer thing?

I believe it passes this test as it’s a reasoning model

Fast open inference makes natural voice UX viable without vendor lock-in.
