Загрузка видео...
Не удалось загрузить видео
I taught a speech model to understand context in conversation. This is what happened It adjusts voice and tone to express urgency, comfort, understanding from the dialogue. Just like a real human being 520M model. Runs locally on consumer devices How this is achieved 🧵
22,028 просмотров • 6 месяцев назад •via X (Twitter)
Комментарии: 28

Full report:

How do we know someone understands us? Mostly by their voice, rarely by words alone But most speech models today are still text → audio, and they drop the context behind the voice So AI voices sound mechanical, rather than like a companion who enjoys conversation

We make it practical with a neural audio codec and a hierarchical Transformer Big backbone. Small decoder. This saves ~97% of the heavy compute Workflow (video): 1. Feed text + speech tokens into the backbone in parallel 2. The backbone summarizes everything into one context vector 3. A tiny Depformer autoregressively predicts codec tokens 4. The codec decodes tokens into speech

We unlocked something bigger: contextual adaptation No human labels. It learns tone from context How? Traditional TTS only sees the sentence, so it reads “I’m fine” the same every time We feed it the full conversation history Now the voice matches the moment Same text. Different emotion.

It runs fully on-device on consumer hardware - NVIDIA RTX 30/40/50 series - Apple Silicon M1–M4 Siri could use a voice like this too

What are we building with this voice? The next generation of games at @EntropyGamesAI No dialogue trees No pre-recorded voice lines NPCs understand player intent, then adjust tone and actions to match the moment

People say “gamers hate AI.” Not true. Gamers hate fake AI If you build a great game experience, players will tell you: 1k upvotes, 400+ comments on our AI companion NPC prototype in the first 24 hours This is what demand pull looks like

Voice is the most human interface When a voice understands context, it stops sounding like a model It starts sounding like someone We’re building a future where every character and assistant can hear the moment, not just the words

this is cool also appreciate the local-first approach when many are doing the opposite

yeahhh 🫡 Beyond the economic costs, a direct benefit of local-first is that it alleviates user privacy concerns by ensuring all data remains stored on the local device. You might not worry about AI knowing the data from your code, but you certainly would worry about AI knowing the data from your most intimate conversations.

For 50 years, game characters had bodies and worlds built around them. But their voice was a recording. Read once, replayed forever. This changes that. A voice that hears the scene and responds in the moment. 520M on-device speech model, running locally on consumer GPUs at zero inference cost. Proud of the team for shipping this!

Just ship it! :D

Love this, when can I play this in my project zomboid

Glad you like it! Do you have any specific thoughts or requests regarding AI NPCs? Especially coming from someone who's played thousands of hours across hundreds of games

I’ve been there watching this journey, so proud of u and @ChrisYicheng !

@ChrisYicheng Thanks! Are you still gaming? Or mostly vibe coding lol

this is scarily cool. Well done

wow! awesome!

Very cool - gg

Thanks!

kinda reminds me of @sesame

very cool!

Thank you!!

Cool. We need to chat.

👀

this is really just smart use of context and prosody control

this is amazing 👍

Thank you!
