Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Speech-to-speech no longer needs speech-to-text! Until now, our stack was VAD -> STT -> LLM -> TTS. Now it can send audio directly to multimodal LLMs: VAD → MLLM → TTS No STT. The model understands your voice. Now go build better voice agents!

119,372 görüntüleme • 1 ay önce •via X (Twitter)

74 Yorum

Julien Blanchon 🇺🇦 profil fotoğrafı
Julien Blanchon 🇺🇦1 ay önce

It's pretty old but you can even MLLMify any LLM with

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

I didn't remember, very cool!

moskstraumen profil fotoğrafı
moskstraumen1 ay önce

Ok but.. Apart from gemma 4 e2b/e4b, no other model (afaik) understands audio directly.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

There are a bunch! inkling, phi-multimodal, qwen2-audio, ...

moskstraumen profil fotoğrafı
moskstraumen1 ay önce

True, I forgot inkling (among the usable ones).

Diego Carlino profil fotoğrafı
Diego Carlino1 ay önce

we are only missing full duplex now 🔥

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

but do we want it?

Jon Hulsinger profil fotoğrafı
Jon Hulsinger1 ay önce

@_carlid_dev Yes

Victor Escorcia profil fotoğrafı
Victor Escorcia1 ay önce

@andimarafioti @_carlid_dev Exactly, why not? The cultural norms about talking to AI aren't written yet. Actually, Isn't it that only focusing on half-duplex & semiduplex mostly laziness from an ML design standpoint & blocks new UX & ML designs/future?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

@jhulsinger @_carlid_dev I think it's mostly about a lack of data. You can train the models, but they end up not being that good at it. Lack of data also means you can't scale them up that much.

Carson profil fotoğrafı
Carson1 ay önce

@3scorciav @jhulsinger @_carlid_dev Full duplex is tough because of how you have to balance latency and quality. If you want the real benefits (back channeling and barge-in), then you need fast auto-regression that doesn’t suck. We aren’t there quite yet

Victor Escorcia profil fotoğrafı
Victor Escorcia1 ay önce

@andimarafioti @jhulsinger @_carlid_dev That's a technical thing! What kind of task & benchmark should researchers (academic or not) use/create?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

@twinpeakstownie @jhulsinger @_carlid_dev Several people building kids toys with this :)

Aadeel Khaan profil fotoğrafı
Aadeel Khaan1 ay önce

Speech-to-speech is a real unlock, and the place it'll get tested hardest is telephony. Phone audio is 8kHz narrowband with codec artifacts, and most multimodal audio models are trained and evaluated on clean wideband. That gap is where the demo-to-production drop lives. The other thing that keeps STT in our stack: we need the transcript, not just the response. The booking record, the allergy the guest mentioned, what was actually agreed on the call. A model that understands the audio but doesn't hand back text solves the conversation and loses the record. Not an argument against the architecture. Just the two constraints that decide whether we can drop a stage.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

That makes sense! And I can look into phone calls next 🙌

Aadeel Khaan profil fotoğrafı
Aadeel Khaan1 ay önce

That'd be a genuinely useful eval, and if you want the hard mode: 8kHz narrowband, G.711 codec, dining-room noise behind the caller, and people correcting themselves mid-sentence ("table for four, actually five"). That's where the public corpora stop resembling reality.

Aadeel Khaan profil fotoğrafı
Aadeel Khaan1 ay önce

We open-sourced the harness we built for this: sttbench, replays your own recordings against providers at real time and reports WER, finalization lag p50/p95, and cost per hour side by side. npx sttbench run ./calls, MIT, your keys. Might save you the scaffolding if you go down the phone-call path.

pulsatingGenius profil fotoğrafı
pulsatingGenius1 ay önce

Isn’t there any degradation in intelligence/reasoning when speech is directly used without any inner monologue of text?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

It all depends on the model. For some, it's the opposite (they reason better about audio). But in general, my impression is that they are better with text inputs, there's just more data. In any case, models can get input audio and reason with text.

Alec Freudenstein profil fotoğrafı
Alec Freudenstein1 ay önce

Insane work. Well done sir! 🎩👏

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Thank you!

Craig Krull day190 profil fotoğrafı
Craig Krull day1901 ay önce

That’s really amazing , but my word that’s ( for me ) an irritating ai voice

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

voice cloning is pretty good with this TTS :)

Dev Patel profil fotoğrafı
Dev Patel1 ay önce

@grok how does this work and how does it compare to traditional s2s

Rohit profil fotoğrafı
Rohit1 ay önce

Neat collapse of the stack. Interesting counter-case: for dictation into apps, STT is still the product — you want formatted text at the cursor, instantly. That's the layer we build at @helloattyn. Direct audio→MLLM makes total sense for voice agents though.

Amr Kayid 🪼 profil fotoğrafı
Amr Kayid 🪼1 ay önce

When will we be able to replace VAD too?

Luis Mata profil fotoğrafı
Luis Mata1 ay önce

This is amazing

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Thank you!

Manuel Martinez profil fotoğrafı
Manuel Martinez1 ay önce

@gonzalo_io a bit long the lines of what we were talking about yesterday!

Alexa | Indie hacker profil fotoğrafı
Alexa | Indie hacker1 ay önce

really good build.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Thank you!

Vishal Singh 🥑 profil fotoğrafı
Vishal Singh 🥑27 gün önce

damn solid build 🔥

Andi Marafioti profil fotoğrafı
Andi Marafioti27 gün önce

thank you!

StripSlashes profil fotoğrafı
StripSlashes1 ay önce

this is like a month old. keep up

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

yeah, I merged a few other PRs in between but I hadn't announced it yet and realized some people didn't know :)

lmao profil fotoğrafı
lmao1 ay önce

no stt doesn't mean no transcript. you just moved the hard part into a pricier black box

Peter MOUEZA🇲🇫 profil fotoğrafı
Peter MOUEZA🇲🇫1 ay önce

Audio

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan1 ay önce

dropping stt cuts latency, but makes production debugging much harder without searchable transcripts

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

we don't store transcripts in production... pii and all

WikiDB profil fotoğrafı
WikiDB1 ay önce

What about end to end voice translation model?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

you can definitely build that with a simple prompt here :)

WikiDB profil fotoğrafı
WikiDB1 ay önce

Nice! I will definitely try this

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Awesome!

Adel Bucetta profil fotoğrafı
Adel Bucetta1 ay önce

the honest answer is that most people still think stt is the hard part, when really it's just the entry ticket. now we can focus on what actually matters: the model understanding and responding to voice, not just transcribing it.

Alberto Andreotti. El Capitan Beto. profil fotoğrafı
Alberto Andreotti. El Capitan Beto.1 ay önce

how do you come back to a normal life after that performance? 😆

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

life is my stage xD

Mykyta Pavlenko profil fotoğrafı
Mykyta Pavlenko1 ay önce

killing the stt step changes more than latency: you stop losing tone, hesitation, and emphasis in the transcript. that is where the human feel was dying

Glafira Firana profil fotoğrafı
Glafira Firana1 ay önce

ngl curious how this handles accents and background noise compared to the old way. seems like it could go either way?

Suyash profil fotoğrafı
Suyash1 ay önce

Woah this is nice!

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Thank you! It was hard work 🥹

Nenad Mancevic profil fotoğrafı
Nenad Mancevic1 ay önce

Super! Have you done some testing on which of the mllms performs best?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

Yes! Inkling performs best, but it's hard to run it on normal hardware (I used 4x6000s :|). Gemma 4 12B is great, but I think I prefer using STT and 32B

Nenad Mancevic profil fotoğrafı
Nenad Mancevic1 ay önce

Awesome! Would be great to also compare the end WER using the direct mllm in/out vs standard stt, tts approach. If the stt fails then obviously the tts will amplify that. It’s been challenging to get WER below< 5% using the whispers of the world.

DEV profil fotoğrafı
DEV1 ay önce

This shift cuts latency significantly while enhancing user experience. Simple setups just became...

Praveen Kumar Verma profil fotoğrafı
Praveen Kumar Verma1 ay önce

Interesting, do we have guardrails or encryption for filtering PII data for security?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

You can deploy this locally 🙌 and in the deployed version we serve we don’t store anything

Praveen Kumar Verma profil fotoğrafı
Praveen Kumar Verma1 ay önce

Great

Harshit Sharma profil fotoğrafı
Harshit Sharma1 ay önce

lmao

Manu_TechAndGames profil fotoğrafı
Manu_TechAndGames1 ay önce

The GitHub readme says it's using parakeet as a stt. Did I misunderstand something?

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

that's the default setting, you can pass -stt none to bypass any stt, but you then need an llm that has audio input. The readme has info about it, here's an example as well

Manu_TechAndGames profil fotoğrafı
Manu_TechAndGames1 ay önce

Ok, thanks , I missed it !

Jack GM profil fotoğrafı
Jack GM1 ay önce

This simplifies the voice-agent stack a lot. Removing STT as a separate brittle step should preserve tone, timing, and hesitation signals—useful for assistants that need to understand intent, not just words.

Sidhu Ram profil fotoğrafı
Sidhu Ram1 ay önce

VAD-STT-LMM-TTS is helpful in scenarios where we need audit trails of the entire audio. How do voice agents solve the need for proper end to end audit's of the sessions?

Daniel Young 🫰🏻 profil fotoğrafı
Daniel Young 🫰🏻1 ay önce

Am I crazy or something? A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS Is the first line of the readme.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

You’re right! I have to adapt the readme! Will do it today. It already says that you can skip the stt

Daniel Young 🫰🏻 profil fotoğrafı
Daniel Young 🫰🏻1 ay önce

Ah I see it way down at the bottom. What’s the quality of the speech to speech and what latency? Can it really be realtime? Everything speech to speech I’ve done has had to be local to get it under the 50ms human predicable delay.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

50ms is way faster than humans! Try our demo here:

Daniel Young 🫰🏻 profil fotoğrafı
Daniel Young 🫰🏻1 ay önce

Will do! My current investigations have been for real time strategy game applications. 300ms is a VERY noticeable lag. For other seminal time applications it should be fine. Beatrice2 has been REALLY good and fun to fine tune with Gemini TTS to create character voices.

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

How are you measuring it? 300ms is so short, at that latency models feels super annoying to me, they cut you all the time before you're done, and even when not it feels super eager. Humans take longer to reply.

Daniel Young 🫰🏻 profil fotoğrafı
Daniel Young 🫰🏻1 ay önce

Oh! I think we’re talking about two different pipelines lol. For the use case in real time gaming, it’s user speech -> speech to speech model to change their perceived voice -> VOIP -> other users output device. Use case is more for novelty and anonymity vs agentic conversation

Andi Marafioti profil fotoğrafı
Andi Marafioti1 ay önce

ahhh ok, yeah, there the low latency makes more sense :)

Gregor profil fotoğrafı
Gregor1 ay önce

What does the latency look like compared to Whisper small? The MLLM step seems considerably heavier than a dedicated STT model.

AINZX profil fotoğrafı
AINZX1 ay önce

Model size

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

Skipping STT entirely, that's a real shift

Benzer Videolar