Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Speech-to-speech no longer needs speech-to-text! Until now, our stack was VAD -> STT -> LLM -> TTS. Now it can send audio directly to multimodal LLMs: VAD → MLLM → TTS No STT. The model understands your voice. Now go build better voice agents!

119,372 Aufrufe • vor 1 Monat •via X (Twitter)

74 Kommentare

Profilbild von Julien Blanchon 🇺🇦
Julien Blanchon 🇺🇦vor 1 Monat

It's pretty old but you can even MLLMify any LLM with

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

I didn't remember, very cool!

Profilbild von moskstraumen
moskstraumenvor 1 Monat

Ok but.. Apart from gemma 4 e2b/e4b, no other model (afaik) understands audio directly.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

There are a bunch! inkling, phi-multimodal, qwen2-audio, ...

Profilbild von moskstraumen
moskstraumenvor 1 Monat

True, I forgot inkling (among the usable ones).

Profilbild von Diego Carlino
Diego Carlinovor 1 Monat

we are only missing full duplex now 🔥

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

but do we want it?

Profilbild von Jon Hulsinger
Jon Hulsingervor 1 Monat

@_carlid_dev Yes

Profilbild von Victor Escorcia
Victor Escorciavor 1 Monat

@andimarafioti @_carlid_dev Exactly, why not? The cultural norms about talking to AI aren't written yet. Actually, Isn't it that only focusing on half-duplex & semiduplex mostly laziness from an ML design standpoint & blocks new UX & ML designs/future?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

@jhulsinger @_carlid_dev I think it's mostly about a lack of data. You can train the models, but they end up not being that good at it. Lack of data also means you can't scale them up that much.

Profilbild von Carson
Carsonvor 1 Monat

@3scorciav @jhulsinger @_carlid_dev Full duplex is tough because of how you have to balance latency and quality. If you want the real benefits (back channeling and barge-in), then you need fast auto-regression that doesn’t suck. We aren’t there quite yet

Profilbild von Victor Escorcia
Victor Escorciavor 1 Monat

@andimarafioti @jhulsinger @_carlid_dev That's a technical thing! What kind of task & benchmark should researchers (academic or not) use/create?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

@twinpeakstownie @jhulsinger @_carlid_dev Several people building kids toys with this :)

Profilbild von Aadeel Khaan
Aadeel Khaanvor 1 Monat

Speech-to-speech is a real unlock, and the place it'll get tested hardest is telephony. Phone audio is 8kHz narrowband with codec artifacts, and most multimodal audio models are trained and evaluated on clean wideband. That gap is where the demo-to-production drop lives. The other thing that keeps STT in our stack: we need the transcript, not just the response. The booking record, the allergy the guest mentioned, what was actually agreed on the call. A model that understands the audio but doesn't hand back text solves the conversation and loses the record. Not an argument against the architecture. Just the two constraints that decide whether we can drop a stage.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

That makes sense! And I can look into phone calls next 🙌

Profilbild von Aadeel Khaan
Aadeel Khaanvor 1 Monat

That'd be a genuinely useful eval, and if you want the hard mode: 8kHz narrowband, G.711 codec, dining-room noise behind the caller, and people correcting themselves mid-sentence ("table for four, actually five"). That's where the public corpora stop resembling reality.

Profilbild von Aadeel Khaan
Aadeel Khaanvor 1 Monat

We open-sourced the harness we built for this: sttbench, replays your own recordings against providers at real time and reports WER, finalization lag p50/p95, and cost per hour side by side. npx sttbench run ./calls, MIT, your keys. Might save you the scaffolding if you go down the phone-call path.

Profilbild von pulsatingGenius
pulsatingGeniusvor 1 Monat

Isn’t there any degradation in intelligence/reasoning when speech is directly used without any inner monologue of text?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

It all depends on the model. For some, it's the opposite (they reason better about audio). But in general, my impression is that they are better with text inputs, there's just more data. In any case, models can get input audio and reason with text.

Profilbild von Alec Freudenstein
Alec Freudensteinvor 1 Monat

Insane work. Well done sir! 🎩👏

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Thank you!

Profilbild von Craig Krull day190
Craig Krull day190vor 1 Monat

That’s really amazing , but my word that’s ( for me ) an irritating ai voice

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

voice cloning is pretty good with this TTS :)

Profilbild von Dev Patel
Dev Patelvor 1 Monat

@grok how does this work and how does it compare to traditional s2s

Profilbild von Rohit
Rohitvor 1 Monat

Neat collapse of the stack. Interesting counter-case: for dictation into apps, STT is still the product — you want formatted text at the cursor, instantly. That's the layer we build at @helloattyn. Direct audio→MLLM makes total sense for voice agents though.

Profilbild von Amr Kayid 🪼
Amr Kayid 🪼vor 1 Monat

When will we be able to replace VAD too?

Profilbild von Luis Mata
Luis Matavor 1 Monat

This is amazing

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Thank you!

Profilbild von Manuel Martinez
Manuel Martinezvor 1 Monat

@gonzalo_io a bit long the lines of what we were talking about yesterday!

Profilbild von Alexa | Indie hacker
Alexa | Indie hackervor 1 Monat

really good build.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Thank you!

Profilbild von Vishal Singh 🥑
Vishal Singh 🥑vor 27 Tagen

damn solid build 🔥

Profilbild von Andi Marafioti
Andi Marafiotivor 27 Tagen

thank you!

Profilbild von StripSlashes
StripSlashesvor 1 Monat

this is like a month old. keep up

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

yeah, I merged a few other PRs in between but I hadn't announced it yet and realized some people didn't know :)

Profilbild von lmao
lmaovor 1 Monat

no stt doesn't mean no transcript. you just moved the hard part into a pricier black box

Profilbild von Peter MOUEZA🇲🇫
Peter MOUEZA🇲🇫vor 1 Monat

Audio

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 1 Monat

dropping stt cuts latency, but makes production debugging much harder without searchable transcripts

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

we don't store transcripts in production... pii and all

Profilbild von WikiDB
WikiDBvor 1 Monat

What about end to end voice translation model?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

you can definitely build that with a simple prompt here :)

Profilbild von WikiDB
WikiDBvor 1 Monat

Nice! I will definitely try this

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Awesome!

Profilbild von Adel Bucetta
Adel Bucettavor 1 Monat

the honest answer is that most people still think stt is the hard part, when really it's just the entry ticket. now we can focus on what actually matters: the model understanding and responding to voice, not just transcribing it.

Profilbild von Alberto Andreotti. El Capitan Beto.
Alberto Andreotti. El Capitan Beto.vor 1 Monat

how do you come back to a normal life after that performance? 😆

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

life is my stage xD

Profilbild von Mykyta Pavlenko
Mykyta Pavlenkovor 1 Monat

killing the stt step changes more than latency: you stop losing tone, hesitation, and emphasis in the transcript. that is where the human feel was dying

Profilbild von Glafira Firana
Glafira Firanavor 1 Monat

ngl curious how this handles accents and background noise compared to the old way. seems like it could go either way?

Profilbild von Suyash
Suyashvor 1 Monat

Woah this is nice!

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Thank you! It was hard work 🥹

Profilbild von Nenad Mancevic
Nenad Mancevicvor 1 Monat

Super! Have you done some testing on which of the mllms performs best?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

Yes! Inkling performs best, but it's hard to run it on normal hardware (I used 4x6000s :|). Gemma 4 12B is great, but I think I prefer using STT and 32B

Profilbild von Nenad Mancevic
Nenad Mancevicvor 1 Monat

Awesome! Would be great to also compare the end WER using the direct mllm in/out vs standard stt, tts approach. If the stt fails then obviously the tts will amplify that. It’s been challenging to get WER below< 5% using the whispers of the world.

Profilbild von DEV
DEVvor 1 Monat

This shift cuts latency significantly while enhancing user experience. Simple setups just became...

Profilbild von Praveen Kumar Verma
Praveen Kumar Vermavor 1 Monat

Interesting, do we have guardrails or encryption for filtering PII data for security?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

You can deploy this locally 🙌 and in the deployed version we serve we don’t store anything

Profilbild von Praveen Kumar Verma
Praveen Kumar Vermavor 1 Monat

Great

Profilbild von Harshit Sharma
Harshit Sharmavor 1 Monat

lmao

Profilbild von Manu_TechAndGames
Manu_TechAndGamesvor 1 Monat

The GitHub readme says it's using parakeet as a stt. Did I misunderstand something?

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

that's the default setting, you can pass -stt none to bypass any stt, but you then need an llm that has audio input. The readme has info about it, here's an example as well

Profilbild von Manu_TechAndGames
Manu_TechAndGamesvor 1 Monat

Ok, thanks , I missed it !

Profilbild von Jack GM
Jack GMvor 1 Monat

This simplifies the voice-agent stack a lot. Removing STT as a separate brittle step should preserve tone, timing, and hesitation signals—useful for assistants that need to understand intent, not just words.

Profilbild von Sidhu Ram
Sidhu Ramvor 1 Monat

VAD-STT-LMM-TTS is helpful in scenarios where we need audit trails of the entire audio. How do voice agents solve the need for proper end to end audit's of the sessions?

Profilbild von Daniel Young 🫰🏻
Daniel Young 🫰🏻vor 1 Monat

Am I crazy or something? A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS Is the first line of the readme.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

You’re right! I have to adapt the readme! Will do it today. It already says that you can skip the stt

Profilbild von Daniel Young 🫰🏻
Daniel Young 🫰🏻vor 1 Monat

Ah I see it way down at the bottom. What’s the quality of the speech to speech and what latency? Can it really be realtime? Everything speech to speech I’ve done has had to be local to get it under the 50ms human predicable delay.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

50ms is way faster than humans! Try our demo here:

Profilbild von Daniel Young 🫰🏻
Daniel Young 🫰🏻vor 1 Monat

Will do! My current investigations have been for real time strategy game applications. 300ms is a VERY noticeable lag. For other seminal time applications it should be fine. Beatrice2 has been REALLY good and fun to fine tune with Gemini TTS to create character voices.

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

How are you measuring it? 300ms is so short, at that latency models feels super annoying to me, they cut you all the time before you're done, and even when not it feels super eager. Humans take longer to reply.

Profilbild von Daniel Young 🫰🏻
Daniel Young 🫰🏻vor 1 Monat

Oh! I think we’re talking about two different pipelines lol. For the use case in real time gaming, it’s user speech -> speech to speech model to change their perceived voice -> VOIP -> other users output device. Use case is more for novelty and anonymity vs agentic conversation

Profilbild von Andi Marafioti
Andi Marafiotivor 1 Monat

ahhh ok, yeah, there the low latency makes more sense :)

Profilbild von Gregor
Gregorvor 1 Monat

What does the latency look like compared to Whisper small? The MLLM step seems considerably heavier than a dedicated STT model.

Profilbild von AINZX
AINZXvor 1 Monat

Model size

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Skipping STT entirely, that's a real shift

Ähnliche Videos