Video wird geladen...
Video konnte nicht geladen werden
Speech-to-speech no longer needs speech-to-text! Until now, our stack was VAD -> STT -> LLM -> TTS. Now it can send audio directly to multimodal LLMs: VAD → MLLM → TTS No STT. The model understands your voice. Now go build better voice agents!
119,372 Aufrufe • vor 1 Monat •via X (Twitter)
74 Kommentare

It's pretty old but you can even MLLMify any LLM with

I didn't remember, very cool!

Ok but.. Apart from gemma 4 e2b/e4b, no other model (afaik) understands audio directly.

There are a bunch! inkling, phi-multimodal, qwen2-audio, ...

True, I forgot inkling (among the usable ones).

we are only missing full duplex now 🔥

but do we want it?

@_carlid_dev Yes

@andimarafioti @_carlid_dev Exactly, why not? The cultural norms about talking to AI aren't written yet. Actually, Isn't it that only focusing on half-duplex & semiduplex mostly laziness from an ML design standpoint & blocks new UX & ML designs/future?

@jhulsinger @_carlid_dev I think it's mostly about a lack of data. You can train the models, but they end up not being that good at it. Lack of data also means you can't scale them up that much.

@3scorciav @jhulsinger @_carlid_dev Full duplex is tough because of how you have to balance latency and quality. If you want the real benefits (back channeling and barge-in), then you need fast auto-regression that doesn’t suck. We aren’t there quite yet

@andimarafioti @jhulsinger @_carlid_dev That's a technical thing! What kind of task & benchmark should researchers (academic or not) use/create?

@twinpeakstownie @jhulsinger @_carlid_dev Several people building kids toys with this :)

Speech-to-speech is a real unlock, and the place it'll get tested hardest is telephony. Phone audio is 8kHz narrowband with codec artifacts, and most multimodal audio models are trained and evaluated on clean wideband. That gap is where the demo-to-production drop lives. The other thing that keeps STT in our stack: we need the transcript, not just the response. The booking record, the allergy the guest mentioned, what was actually agreed on the call. A model that understands the audio but doesn't hand back text solves the conversation and loses the record. Not an argument against the architecture. Just the two constraints that decide whether we can drop a stage.

That makes sense! And I can look into phone calls next 🙌

That'd be a genuinely useful eval, and if you want the hard mode: 8kHz narrowband, G.711 codec, dining-room noise behind the caller, and people correcting themselves mid-sentence ("table for four, actually five"). That's where the public corpora stop resembling reality.

We open-sourced the harness we built for this: sttbench, replays your own recordings against providers at real time and reports WER, finalization lag p50/p95, and cost per hour side by side. npx sttbench run ./calls, MIT, your keys. Might save you the scaffolding if you go down the phone-call path.

Isn’t there any degradation in intelligence/reasoning when speech is directly used without any inner monologue of text?

It all depends on the model. For some, it's the opposite (they reason better about audio). But in general, my impression is that they are better with text inputs, there's just more data. In any case, models can get input audio and reason with text.

Insane work. Well done sir! 🎩👏

Thank you!

That’s really amazing , but my word that’s ( for me ) an irritating ai voice

voice cloning is pretty good with this TTS :)

@grok how does this work and how does it compare to traditional s2s

Neat collapse of the stack. Interesting counter-case: for dictation into apps, STT is still the product — you want formatted text at the cursor, instantly. That's the layer we build at @helloattyn. Direct audio→MLLM makes total sense for voice agents though.

When will we be able to replace VAD too?

This is amazing

Thank you!

@gonzalo_io a bit long the lines of what we were talking about yesterday!

really good build.

Thank you!

damn solid build 🔥

thank you!

this is like a month old. keep up

yeah, I merged a few other PRs in between but I hadn't announced it yet and realized some people didn't know :)

no stt doesn't mean no transcript. you just moved the hard part into a pricier black box

Audio

dropping stt cuts latency, but makes production debugging much harder without searchable transcripts

we don't store transcripts in production... pii and all

What about end to end voice translation model?

you can definitely build that with a simple prompt here :)

Nice! I will definitely try this

Awesome!

the honest answer is that most people still think stt is the hard part, when really it's just the entry ticket. now we can focus on what actually matters: the model understanding and responding to voice, not just transcribing it.

how do you come back to a normal life after that performance? 😆

life is my stage xD

killing the stt step changes more than latency: you stop losing tone, hesitation, and emphasis in the transcript. that is where the human feel was dying

ngl curious how this handles accents and background noise compared to the old way. seems like it could go either way?

Woah this is nice!

Thank you! It was hard work 🥹

Super! Have you done some testing on which of the mllms performs best?

Yes! Inkling performs best, but it's hard to run it on normal hardware (I used 4x6000s :|). Gemma 4 12B is great, but I think I prefer using STT and 32B

Awesome! Would be great to also compare the end WER using the direct mllm in/out vs standard stt, tts approach. If the stt fails then obviously the tts will amplify that. It’s been challenging to get WER below< 5% using the whispers of the world.

This shift cuts latency significantly while enhancing user experience. Simple setups just became...

Interesting, do we have guardrails or encryption for filtering PII data for security?

You can deploy this locally 🙌 and in the deployed version we serve we don’t store anything

Great

lmao

The GitHub readme says it's using parakeet as a stt. Did I misunderstand something?

that's the default setting, you can pass -stt none to bypass any stt, but you then need an llm that has audio input. The readme has info about it, here's an example as well

Ok, thanks , I missed it !

This simplifies the voice-agent stack a lot. Removing STT as a separate brittle step should preserve tone, timing, and hesitation signals—useful for assistants that need to understand intent, not just words.

VAD-STT-LMM-TTS is helpful in scenarios where we need audit trails of the entire audio. How do voice agents solve the need for proper end to end audit's of the sessions?

Am I crazy or something? A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS Is the first line of the readme.

You’re right! I have to adapt the readme! Will do it today. It already says that you can skip the stt

Ah I see it way down at the bottom. What’s the quality of the speech to speech and what latency? Can it really be realtime? Everything speech to speech I’ve done has had to be local to get it under the 50ms human predicable delay.

50ms is way faster than humans! Try our demo here:

Will do! My current investigations have been for real time strategy game applications. 300ms is a VERY noticeable lag. For other seminal time applications it should be fine. Beatrice2 has been REALLY good and fun to fine tune with Gemini TTS to create character voices.

How are you measuring it? 300ms is so short, at that latency models feels super annoying to me, they cut you all the time before you're done, and even when not it feels super eager. Humans take longer to reply.

Oh! I think we’re talking about two different pipelines lol. For the use case in real time gaming, it’s user speech -> speech to speech model to change their perceived voice -> VOIP -> other users output device. Use case is more for novelty and anonymity vs agentic conversation

ahhh ok, yeah, there the low latency makes more sense :)

What does the latency look like compared to Whisper small? The MLLM step seems considerably heavier than a dedicated STT model.

Model size

Skipping STT entirely, that's a real shift
