Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required. Open-source, lightweight, and incredibly fast. 🧵👇

238,618 Aufrufe • vor 8 Monaten •via X (Twitter)

40 Kommentare

Profilbild von kyutai
kyutaivor 8 Monaten

The gap in TTS today: ✘ Huge LLM models (1B+ params) need a GPU. ✘ Small models like Kokoro (82M) are fast but lack flexible voice cloning. Pocket TTS bridges the gap. It runs faster than real-time on an average laptop’s CPU while keeping the power of the giants.

Profilbild von kyutai
kyutaivor 8 Monaten

🎙️ True voice cloning: Pocket TTS only needs 5 seconds of audio to capture: ✔ Individual voice "color" ✔ the emotion and accent ✔ Acoustic conditions (reverb, mic quality) Use our voice library or clone any voice from a tiny sample.

Profilbild von kyutai
kyutaivor 8 Monaten

📊 The stats don't lie. Despite its size (100M params), Pocket TTS beats F5-TTS and DSM in Word Error Rate (1.84) and Audio Quality ELO. It’s the only model in its class that offers voice cloning *and* runs comfortably on a CPU.

Profilbild von kyutai
kyutaivor 8 Monaten

🏗️ How? We moved away from discrete tokens. Pocket TTS is based on the Continuous Audio Language Models (CALM) framework and predicts sequences of continuous latents directly, using a 1-step sampling method (Lagrangian Self-Distillation). CALM paper:

Profilbild von kyutai
kyutaivor 8 Monaten

🔓 Open source for everyone. Trained on 88k hours of public English data to ensure reproducibility. Check out the code and the technical breakdown here:

Profilbild von Raj Breno
Raj Brenovor 8 Monaten

Is it multilingual?

Profilbild von kyutai
kyutaivor 8 Monaten

English-only for now.

Profilbild von Steve grins
Steve grinsvor 8 Monaten

ran a few samples on my m2 air. startup is instant and voices are scarily close to the original. wild to see this quality without a gpu or cloud. nice work

Profilbild von Seba
Sebavor 8 Monaten

CPU inference is old news

Profilbild von Fahd Mirza
Fahd Mirzavor 8 Monaten

🚨 THIS 100MB TTS MODEL RUNS ON YOUR LAPTOP 🎆 ♠ and it's 6x faster than real-time 🚀 🔹 Pocket-TTS = voice cloning with zero GPU requirements 🔹 CALM architecture skips token generation entirely 🔹 200ms latency + offline + handles infinite text 🔹 How continuous modeling killed expensive APIs 🔥 Watch the full breakdown here:

Profilbild von Ananya Patelik
Ananya Patelikvor 8 Monaten

100M params and no GPU? that’s wizardry. tried it on my old thinkpad and it sings. voice cloning is scarily good. hats off, team!

Profilbild von Amos Gyamfi
Amos Gyamfivor 8 Monaten

Integrating this with will be awwwesome. I'm eager to try it.

Profilbild von SebastianBoo
SebastianBoovor 8 Monaten

Can you share the train code? 🥺

Profilbild von jezza kezza
jezza kezzavor 8 Monaten

I use voice clone TTS every day. Fine as an 'AI character' model, well done. Not good enough for things like podcasts, dubbing or back-patching existing recording as it loses some natural character. To be fair, even elevenLabs is borderline for that in professional settings, so no shade for a tiny model ! Maybe OK for promo uses.

Profilbild von Dark
Darkvor 8 Monaten

need to inline it with PDF controls. 95% of the TTS I use is reading PDFs. read highlighted section, page, etc

Profilbild von jim
jimvor 8 Monaten

Got it running in a Docker container and working as my Home Assistant Voice. Much nicer results than Piper and about the same speed!!! Very good

Profilbild von kaiyes
kaiyesvor 8 Monaten

would love to use this in mobile. Any implementation yet ? Also, can it do Dual language at the same time ? Like 2 languages spoken in the same sentence - can it catch that ?

Profilbild von Matthew Man e/acc ⏩
Matthew Man e/acc ⏩vor 8 Monaten

How does it perform next to VibeVoice? Having a benchmark on your project page would be more legit. There’s so many projects, so when something ships without benchmarks it gets pushed to the “maybe I’ll check this out later” pile. With benchmarks comes priority. FYI

Profilbild von Texas Matt
Texas Mattvor 8 Monaten

post it on r/OpenSourceAI

Profilbild von nle
nlevor 8 Monaten

For normal TTS use cases (not voice cloning), how realistic the sound is comparing to Kokoro? Thanks.

Profilbild von Jim Faster
Jim Fastervor 8 Monaten

Everything that Kyutai does is top tier.

Profilbild von CrossProduct
CrossProductvor 8 Monaten

Whats time to first audio though ?

Profilbild von IvannP
IvannPvor 8 Monaten

Nice French accent :) « On est Français oui oui baguette »

Profilbild von Frmn
Frmnvor 8 Monaten

vibecoded the web UI using antigravity. really nice, i'm able to generate a bit slower using old thinkpad with i5-1145G7. but with download functions, the audio is clear and good.

Profilbild von Chaitany
Chaitanyvor 8 Monaten

This is cool definitely going to try.

Profilbild von Dan
Danvor 8 Monaten

When you said said MacBook Air, did you mean intel or apple silicon? & if the latter, which M series chip? I’m trying to gauge performance

Profilbild von b90
b90vor 8 Monaten

c’est magnifique!

Profilbild von Hamza
Hamzavor 8 Monaten

@1million_ae

Profilbild von Goblin Town Citizen, CFA --rtrd/acc
Goblin Town Citizen, CFA --rtrd/accvor 8 Monaten

gpt sovits seems better

Profilbild von ryu
ryuvor 8 Monaten

Have you tried a short context transformer with k=1, i.e. just a skip connection? Makes me think of JiT with the x0 pred v-loss

Profilbild von Zwecharki of Immense AGI fundamentalism
Zwecharki of Immense AGI fundamentalismvor 8 Monaten

@LocallyAIApp wohoo?

Profilbild von Odartei Isaiah
Odartei Isaiahvor 8 Monaten

@grok can this view my ide and help me learn and also learn along with me. How best can it be used

Profilbild von R1K-42
R1K-42vor 8 Monaten

Which languages are planned to be supported? Do you plan to make the code available for fine-tuning?

Profilbild von PanMan
PanManvor 8 Monaten

@amuldotexe

Profilbild von Idrissi
Idrissivor 8 Monaten

Congrats 🎆

Profilbild von FoundationModels
FoundationModelsvor 8 Monaten

wow just 100M parameters - great training

Profilbild von HappyVolley
HappyVolleyvor 8 Monaten

I can't tell you how hard it is to find a good TTS model. Popular ones get silod or paywalled. We know the ones. I've made some real demonic garbled speech before.....onwards kyutai

Profilbild von #INTERNETofAGENTS
#INTERNETofAGENTSvor 8 Monaten

Amazing. 💻

Profilbild von RelakkesYang
RelakkesYangvor 8 Monaten

cool

Profilbild von Giacomo Arienti
Giacomo Arientivor 8 Monaten

Waiting for a multilingual version, please consider to include Italian 🇮🇹

Ähnliche Videos