Video yükleniyor...
Video Yüklenemedi
We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required. Open-source, lightweight, and incredibly fast. 🧵👇
238,618 görüntüleme • 8 ay önce •via X (Twitter)
40 Yorum

The gap in TTS today: ✘ Huge LLM models (1B+ params) need a GPU. ✘ Small models like Kokoro (82M) are fast but lack flexible voice cloning. Pocket TTS bridges the gap. It runs faster than real-time on an average laptop’s CPU while keeping the power of the giants.

🎙️ True voice cloning: Pocket TTS only needs 5 seconds of audio to capture: ✔ Individual voice "color" ✔ the emotion and accent ✔ Acoustic conditions (reverb, mic quality) Use our voice library or clone any voice from a tiny sample.

📊 The stats don't lie. Despite its size (100M params), Pocket TTS beats F5-TTS and DSM in Word Error Rate (1.84) and Audio Quality ELO. It’s the only model in its class that offers voice cloning *and* runs comfortably on a CPU.

🏗️ How? We moved away from discrete tokens. Pocket TTS is based on the Continuous Audio Language Models (CALM) framework and predicts sequences of continuous latents directly, using a 1-step sampling method (Lagrangian Self-Distillation). CALM paper:

🔓 Open source for everyone. Trained on 88k hours of public English data to ensure reproducibility. Check out the code and the technical breakdown here:

Is it multilingual?

English-only for now.

ran a few samples on my m2 air. startup is instant and voices are scarily close to the original. wild to see this quality without a gpu or cloud. nice work

CPU inference is old news

🚨 THIS 100MB TTS MODEL RUNS ON YOUR LAPTOP 🎆 ♠ and it's 6x faster than real-time 🚀 🔹 Pocket-TTS = voice cloning with zero GPU requirements 🔹 CALM architecture skips token generation entirely 🔹 200ms latency + offline + handles infinite text 🔹 How continuous modeling killed expensive APIs 🔥 Watch the full breakdown here:

100M params and no GPU? that’s wizardry. tried it on my old thinkpad and it sings. voice cloning is scarily good. hats off, team!

Integrating this with will be awwwesome. I'm eager to try it.

Can you share the train code? 🥺

I use voice clone TTS every day. Fine as an 'AI character' model, well done. Not good enough for things like podcasts, dubbing or back-patching existing recording as it loses some natural character. To be fair, even elevenLabs is borderline for that in professional settings, so no shade for a tiny model ! Maybe OK for promo uses.

need to inline it with PDF controls. 95% of the TTS I use is reading PDFs. read highlighted section, page, etc

Got it running in a Docker container and working as my Home Assistant Voice. Much nicer results than Piper and about the same speed!!! Very good

would love to use this in mobile. Any implementation yet ? Also, can it do Dual language at the same time ? Like 2 languages spoken in the same sentence - can it catch that ?

How does it perform next to VibeVoice? Having a benchmark on your project page would be more legit. There’s so many projects, so when something ships without benchmarks it gets pushed to the “maybe I’ll check this out later” pile. With benchmarks comes priority. FYI

post it on r/OpenSourceAI

For normal TTS use cases (not voice cloning), how realistic the sound is comparing to Kokoro? Thanks.

Everything that Kyutai does is top tier.

Whats time to first audio though ?

Nice French accent :) « On est Français oui oui baguette »

vibecoded the web UI using antigravity. really nice, i'm able to generate a bit slower using old thinkpad with i5-1145G7. but with download functions, the audio is clear and good.

This is cool definitely going to try.

When you said said MacBook Air, did you mean intel or apple silicon? & if the latter, which M series chip? I’m trying to gauge performance

c’est magnifique!

@1million_ae

gpt sovits seems better

Have you tried a short context transformer with k=1, i.e. just a skip connection? Makes me think of JiT with the x0 pred v-loss

@LocallyAIApp wohoo?

@grok can this view my ide and help me learn and also learn along with me. How best can it be used

Which languages are planned to be supported? Do you plan to make the code available for fine-tuning?

@amuldotexe

Congrats 🎆

wow just 100M parameters - great training

I can't tell you how hard it is to find a good TTS model. Popular ones get silod or paywalled. We know the ones. I've made some real demonic garbled speech before.....onwards kyutai

Amazing. 💻

cool

Waiting for a multilingual version, please consider to include Italian 🇮🇹
