Загрузка видео...

Не удалось загрузить видео

На главную

Speech-to-speech no longer needs speech-to-text! Until now, our stack was VAD -> STT -> LLM -> TTS. Now it can send audio directly to multimodal LLMs: VAD → MLLM → TTS No STT. The model understands your voice. Now go build better voice agents!

119,372 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 74

Фото профиля Julien Blanchon 🇺🇦
Julien Blanchon 🇺🇦1 месяц назад

It's pretty old but you can even MLLMify any LLM with

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

I didn't remember, very cool!

Фото профиля moskstraumen
moskstraumen1 месяц назад

Ok but.. Apart from gemma 4 e2b/e4b, no other model (afaik) understands audio directly.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

There are a bunch! inkling, phi-multimodal, qwen2-audio, ...

Фото профиля moskstraumen
moskstraumen1 месяц назад

True, I forgot inkling (among the usable ones).

Фото профиля Diego Carlino
Diego Carlino1 месяц назад

we are only missing full duplex now 🔥

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

but do we want it?

Фото профиля Jon Hulsinger
Jon Hulsinger1 месяц назад

@_carlid_dev Yes

Фото профиля Victor Escorcia
Victor Escorcia1 месяц назад

@andimarafioti @_carlid_dev Exactly, why not? The cultural norms about talking to AI aren't written yet. Actually, Isn't it that only focusing on half-duplex & semiduplex mostly laziness from an ML design standpoint & blocks new UX & ML designs/future?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

@jhulsinger @_carlid_dev I think it's mostly about a lack of data. You can train the models, but they end up not being that good at it. Lack of data also means you can't scale them up that much.

Фото профиля Carson
Carson1 месяц назад

@3scorciav @jhulsinger @_carlid_dev Full duplex is tough because of how you have to balance latency and quality. If you want the real benefits (back channeling and barge-in), then you need fast auto-regression that doesn’t suck. We aren’t there quite yet

Фото профиля Victor Escorcia
Victor Escorcia1 месяц назад

@andimarafioti @jhulsinger @_carlid_dev That's a technical thing! What kind of task & benchmark should researchers (academic or not) use/create?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

@twinpeakstownie @jhulsinger @_carlid_dev Several people building kids toys with this :)

Фото профиля Aadeel Khaan
Aadeel Khaan1 месяц назад

Speech-to-speech is a real unlock, and the place it'll get tested hardest is telephony. Phone audio is 8kHz narrowband with codec artifacts, and most multimodal audio models are trained and evaluated on clean wideband. That gap is where the demo-to-production drop lives. The other thing that keeps STT in our stack: we need the transcript, not just the response. The booking record, the allergy the guest mentioned, what was actually agreed on the call. A model that understands the audio but doesn't hand back text solves the conversation and loses the record. Not an argument against the architecture. Just the two constraints that decide whether we can drop a stage.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

That makes sense! And I can look into phone calls next 🙌

Фото профиля Aadeel Khaan
Aadeel Khaan1 месяц назад

That'd be a genuinely useful eval, and if you want the hard mode: 8kHz narrowband, G.711 codec, dining-room noise behind the caller, and people correcting themselves mid-sentence ("table for four, actually five"). That's where the public corpora stop resembling reality.

Фото профиля Aadeel Khaan
Aadeel Khaan1 месяц назад

We open-sourced the harness we built for this: sttbench, replays your own recordings against providers at real time and reports WER, finalization lag p50/p95, and cost per hour side by side. npx sttbench run ./calls, MIT, your keys. Might save you the scaffolding if you go down the phone-call path.

Фото профиля pulsatingGenius
pulsatingGenius1 месяц назад

Isn’t there any degradation in intelligence/reasoning when speech is directly used without any inner monologue of text?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

It all depends on the model. For some, it's the opposite (they reason better about audio). But in general, my impression is that they are better with text inputs, there's just more data. In any case, models can get input audio and reason with text.

Фото профиля Alec Freudenstein
Alec Freudenstein1 месяц назад

Insane work. Well done sir! 🎩👏

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Thank you!

Фото профиля Craig Krull day190
Craig Krull day1901 месяц назад

That’s really amazing , but my word that’s ( for me ) an irritating ai voice

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

voice cloning is pretty good with this TTS :)

Фото профиля Dev Patel
Dev Patel1 месяц назад

@grok how does this work and how does it compare to traditional s2s

Фото профиля Rohit
Rohit1 месяц назад

Neat collapse of the stack. Interesting counter-case: for dictation into apps, STT is still the product — you want formatted text at the cursor, instantly. That's the layer we build at @helloattyn. Direct audio→MLLM makes total sense for voice agents though.

Фото профиля Amr Kayid 🪼
Amr Kayid 🪼1 месяц назад

When will we be able to replace VAD too?

Фото профиля Luis Mata
Luis Mata1 месяц назад

This is amazing

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Thank you!

Фото профиля Manuel Martinez
Manuel Martinez1 месяц назад

@gonzalo_io a bit long the lines of what we were talking about yesterday!

Фото профиля Alexa | Indie hacker
Alexa | Indie hacker1 месяц назад

really good build.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Thank you!

Фото профиля Vishal Singh 🥑
Vishal Singh 🥑27 дней назад

damn solid build 🔥

Фото профиля Andi Marafioti
Andi Marafioti27 дней назад

thank you!

Фото профиля StripSlashes
StripSlashes1 месяц назад

this is like a month old. keep up

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

yeah, I merged a few other PRs in between but I hadn't announced it yet and realized some people didn't know :)

Фото профиля lmao
lmao1 месяц назад

no stt doesn't mean no transcript. you just moved the hard part into a pricier black box

Фото профиля Peter MOUEZA🇲🇫
Peter MOUEZA🇲🇫1 месяц назад

Audio

Фото профиля Sebastian Buzdugan
Sebastian Buzdugan1 месяц назад

dropping stt cuts latency, but makes production debugging much harder without searchable transcripts

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

we don't store transcripts in production... pii and all

Фото профиля WikiDB
WikiDB1 месяц назад

What about end to end voice translation model?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

you can definitely build that with a simple prompt here :)

Фото профиля WikiDB
WikiDB1 месяц назад

Nice! I will definitely try this

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Awesome!

Фото профиля Adel Bucetta
Adel Bucetta1 месяц назад

the honest answer is that most people still think stt is the hard part, when really it's just the entry ticket. now we can focus on what actually matters: the model understanding and responding to voice, not just transcribing it.

Фото профиля Alberto Andreotti. El Capitan Beto.
Alberto Andreotti. El Capitan Beto.1 месяц назад

how do you come back to a normal life after that performance? 😆

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

life is my stage xD

Фото профиля Mykyta Pavlenko
Mykyta Pavlenko1 месяц назад

killing the stt step changes more than latency: you stop losing tone, hesitation, and emphasis in the transcript. that is where the human feel was dying

Фото профиля Glafira Firana
Glafira Firana1 месяц назад

ngl curious how this handles accents and background noise compared to the old way. seems like it could go either way?

Фото профиля Suyash
Suyash1 месяц назад

Woah this is nice!

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Thank you! It was hard work 🥹

Фото профиля Nenad Mancevic
Nenad Mancevic1 месяц назад

Super! Have you done some testing on which of the mllms performs best?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

Yes! Inkling performs best, but it's hard to run it on normal hardware (I used 4x6000s :|). Gemma 4 12B is great, but I think I prefer using STT and 32B

Фото профиля Nenad Mancevic
Nenad Mancevic1 месяц назад

Awesome! Would be great to also compare the end WER using the direct mllm in/out vs standard stt, tts approach. If the stt fails then obviously the tts will amplify that. It’s been challenging to get WER below< 5% using the whispers of the world.

Фото профиля DEV
DEV1 месяц назад

This shift cuts latency significantly while enhancing user experience. Simple setups just became...

Фото профиля Praveen Kumar Verma
Praveen Kumar Verma1 месяц назад

Interesting, do we have guardrails or encryption for filtering PII data for security?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

You can deploy this locally 🙌 and in the deployed version we serve we don’t store anything

Фото профиля Praveen Kumar Verma
Praveen Kumar Verma1 месяц назад

Great

Фото профиля Harshit Sharma
Harshit Sharma1 месяц назад

lmao

Фото профиля Manu_TechAndGames
Manu_TechAndGames1 месяц назад

The GitHub readme says it's using parakeet as a stt. Did I misunderstand something?

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

that's the default setting, you can pass -stt none to bypass any stt, but you then need an llm that has audio input. The readme has info about it, here's an example as well

Фото профиля Manu_TechAndGames
Manu_TechAndGames1 месяц назад

Ok, thanks , I missed it !

Фото профиля Jack GM
Jack GM1 месяц назад

This simplifies the voice-agent stack a lot. Removing STT as a separate brittle step should preserve tone, timing, and hesitation signals—useful for assistants that need to understand intent, not just words.

Фото профиля Sidhu Ram
Sidhu Ram1 месяц назад

VAD-STT-LMM-TTS is helpful in scenarios where we need audit trails of the entire audio. How do voice agents solve the need for proper end to end audit's of the sessions?

Фото профиля Daniel Young 🫰🏻
Daniel Young 🫰🏻1 месяц назад

Am I crazy or something? A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS Is the first line of the readme.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

You’re right! I have to adapt the readme! Will do it today. It already says that you can skip the stt

Фото профиля Daniel Young 🫰🏻
Daniel Young 🫰🏻1 месяц назад

Ah I see it way down at the bottom. What’s the quality of the speech to speech and what latency? Can it really be realtime? Everything speech to speech I’ve done has had to be local to get it under the 50ms human predicable delay.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

50ms is way faster than humans! Try our demo here:

Фото профиля Daniel Young 🫰🏻
Daniel Young 🫰🏻1 месяц назад

Will do! My current investigations have been for real time strategy game applications. 300ms is a VERY noticeable lag. For other seminal time applications it should be fine. Beatrice2 has been REALLY good and fun to fine tune with Gemini TTS to create character voices.

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

How are you measuring it? 300ms is so short, at that latency models feels super annoying to me, they cut you all the time before you're done, and even when not it feels super eager. Humans take longer to reply.

Фото профиля Daniel Young 🫰🏻
Daniel Young 🫰🏻1 месяц назад

Oh! I think we’re talking about two different pipelines lol. For the use case in real time gaming, it’s user speech -> speech to speech model to change their perceived voice -> VOIP -> other users output device. Use case is more for novelty and anonymity vs agentic conversation

Фото профиля Andi Marafioti
Andi Marafioti1 месяц назад

ahhh ok, yeah, there the low latency makes more sense :)

Фото профиля Gregor
Gregor1 месяц назад

What does the latency look like compared to Whisper small? The MLLM step seems considerably heavier than a dedicated STT model.

Фото профиля AINZX
AINZX1 месяц назад

Model size

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

Skipping STT entirely, that's a real shift

Похожие видео