Loading video...

Video Failed to Load

Go Home

We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required. Open-source, lightweight, and incredibly fast. 🧵👇

238,618 views • 8 months ago •via X (Twitter)

40 Comments

kyutai's profile picture
kyutai8 months ago

The gap in TTS today: ✘ Huge LLM models (1B+ params) need a GPU. ✘ Small models like Kokoro (82M) are fast but lack flexible voice cloning. Pocket TTS bridges the gap. It runs faster than real-time on an average laptop’s CPU while keeping the power of the giants.

kyutai's profile picture
kyutai8 months ago

🎙️ True voice cloning: Pocket TTS only needs 5 seconds of audio to capture: ✔ Individual voice "color" ✔ the emotion and accent ✔ Acoustic conditions (reverb, mic quality) Use our voice library or clone any voice from a tiny sample.

kyutai's profile picture
kyutai8 months ago

📊 The stats don't lie. Despite its size (100M params), Pocket TTS beats F5-TTS and DSM in Word Error Rate (1.84) and Audio Quality ELO. It’s the only model in its class that offers voice cloning *and* runs comfortably on a CPU.

kyutai's profile picture
kyutai8 months ago

🏗️ How? We moved away from discrete tokens. Pocket TTS is based on the Continuous Audio Language Models (CALM) framework and predicts sequences of continuous latents directly, using a 1-step sampling method (Lagrangian Self-Distillation). CALM paper:

kyutai's profile picture
kyutai8 months ago

🔓 Open source for everyone. Trained on 88k hours of public English data to ensure reproducibility. Check out the code and the technical breakdown here:

Raj Breno's profile picture
Raj Breno8 months ago

Is it multilingual?

kyutai's profile picture
kyutai8 months ago

English-only for now.

Steve grins's profile picture
Steve grins8 months ago

ran a few samples on my m2 air. startup is instant and voices are scarily close to the original. wild to see this quality without a gpu or cloud. nice work

Seba's profile picture
Seba8 months ago

CPU inference is old news

Fahd Mirza's profile picture
Fahd Mirza8 months ago

🚨 THIS 100MB TTS MODEL RUNS ON YOUR LAPTOP 🎆 ♠ and it's 6x faster than real-time 🚀 🔹 Pocket-TTS = voice cloning with zero GPU requirements 🔹 CALM architecture skips token generation entirely 🔹 200ms latency + offline + handles infinite text 🔹 How continuous modeling killed expensive APIs 🔥 Watch the full breakdown here:

Ananya Patelik's profile picture
Ananya Patelik8 months ago

100M params and no GPU? that’s wizardry. tried it on my old thinkpad and it sings. voice cloning is scarily good. hats off, team!

Amos Gyamfi's profile picture
Amos Gyamfi8 months ago

Integrating this with will be awwwesome. I'm eager to try it.

SebastianBoo's profile picture
SebastianBoo8 months ago

Can you share the train code? 🥺

jezza kezza's profile picture
jezza kezza8 months ago

I use voice clone TTS every day. Fine as an 'AI character' model, well done. Not good enough for things like podcasts, dubbing or back-patching existing recording as it loses some natural character. To be fair, even elevenLabs is borderline for that in professional settings, so no shade for a tiny model ! Maybe OK for promo uses.

Dark's profile picture
Dark8 months ago

need to inline it with PDF controls. 95% of the TTS I use is reading PDFs. read highlighted section, page, etc

jim's profile picture
jim8 months ago

Got it running in a Docker container and working as my Home Assistant Voice. Much nicer results than Piper and about the same speed!!! Very good

kaiyes's profile picture
kaiyes8 months ago

would love to use this in mobile. Any implementation yet ? Also, can it do Dual language at the same time ? Like 2 languages spoken in the same sentence - can it catch that ?

Matthew Man e/acc ⏩'s profile picture
Matthew Man e/acc ⏩8 months ago

How does it perform next to VibeVoice? Having a benchmark on your project page would be more legit. There’s so many projects, so when something ships without benchmarks it gets pushed to the “maybe I’ll check this out later” pile. With benchmarks comes priority. FYI

Texas Matt's profile picture
Texas Matt8 months ago

post it on r/OpenSourceAI

nle's profile picture
nle8 months ago

For normal TTS use cases (not voice cloning), how realistic the sound is comparing to Kokoro? Thanks.

Jim Faster's profile picture
Jim Faster8 months ago

Everything that Kyutai does is top tier.

CrossProduct's profile picture
CrossProduct8 months ago

Whats time to first audio though ?

IvannP's profile picture
IvannP8 months ago

Nice French accent :) « On est Français oui oui baguette »

Frmn's profile picture
Frmn8 months ago

vibecoded the web UI using antigravity. really nice, i'm able to generate a bit slower using old thinkpad with i5-1145G7. but with download functions, the audio is clear and good.

Chaitany's profile picture
Chaitany8 months ago

This is cool definitely going to try.

Dan's profile picture
Dan8 months ago

When you said said MacBook Air, did you mean intel or apple silicon? & if the latter, which M series chip? I’m trying to gauge performance

b90's profile picture
b908 months ago

c’est magnifique!

Hamza's profile picture
Hamza8 months ago

@1million_ae

Goblin Town Citizen, CFA --rtrd/acc's profile picture
Goblin Town Citizen, CFA --rtrd/acc8 months ago

gpt sovits seems better

ryu's profile picture
ryu8 months ago

Have you tried a short context transformer with k=1, i.e. just a skip connection? Makes me think of JiT with the x0 pred v-loss

Zwecharki of Immense AGI fundamentalism's profile picture
Zwecharki of Immense AGI fundamentalism8 months ago

@LocallyAIApp wohoo?

Odartei Isaiah's profile picture
Odartei Isaiah8 months ago

@grok can this view my ide and help me learn and also learn along with me. How best can it be used

R1K-42's profile picture
R1K-428 months ago

Which languages are planned to be supported? Do you plan to make the code available for fine-tuning?

PanMan's profile picture
PanMan8 months ago

@amuldotexe

Idrissi's profile picture
Idrissi8 months ago

Congrats 🎆

FoundationModels's profile picture
FoundationModels8 months ago

wow just 100M parameters - great training

HappyVolley's profile picture
HappyVolley8 months ago

I can't tell you how hard it is to find a good TTS model. Popular ones get silod or paywalled. We know the ones. I've made some real demonic garbled speech before.....onwards kyutai

#INTERNETofAGENTS's profile picture
#INTERNETofAGENTS8 months ago

Amazing. 💻

RelakkesYang's profile picture
RelakkesYang8 months ago

cool

Giacomo Arienti's profile picture
Giacomo Arienti8 months ago

Waiting for a multilingual version, please consider to include Italian 🇮🇹

Related Videos