Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required. Open-source, lightweight, and incredibly fast. 🧵👇

238,618 görüntüleme • 8 ay önce •via X (Twitter)

40 Yorum

kyutai profil fotoğrafı
kyutai8 ay önce

The gap in TTS today: ✘ Huge LLM models (1B+ params) need a GPU. ✘ Small models like Kokoro (82M) are fast but lack flexible voice cloning. Pocket TTS bridges the gap. It runs faster than real-time on an average laptop’s CPU while keeping the power of the giants.

kyutai profil fotoğrafı
kyutai8 ay önce

🎙️ True voice cloning: Pocket TTS only needs 5 seconds of audio to capture: ✔ Individual voice "color" ✔ the emotion and accent ✔ Acoustic conditions (reverb, mic quality) Use our voice library or clone any voice from a tiny sample.

kyutai profil fotoğrafı
kyutai8 ay önce

📊 The stats don't lie. Despite its size (100M params), Pocket TTS beats F5-TTS and DSM in Word Error Rate (1.84) and Audio Quality ELO. It’s the only model in its class that offers voice cloning *and* runs comfortably on a CPU.

kyutai profil fotoğrafı
kyutai8 ay önce

🏗️ How? We moved away from discrete tokens. Pocket TTS is based on the Continuous Audio Language Models (CALM) framework and predicts sequences of continuous latents directly, using a 1-step sampling method (Lagrangian Self-Distillation). CALM paper:

kyutai profil fotoğrafı
kyutai8 ay önce

🔓 Open source for everyone. Trained on 88k hours of public English data to ensure reproducibility. Check out the code and the technical breakdown here:

Raj Breno profil fotoğrafı
Raj Breno8 ay önce

Is it multilingual?

kyutai profil fotoğrafı
kyutai8 ay önce

English-only for now.

Steve grins profil fotoğrafı
Steve grins8 ay önce

ran a few samples on my m2 air. startup is instant and voices are scarily close to the original. wild to see this quality without a gpu or cloud. nice work

Seba profil fotoğrafı
Seba8 ay önce

CPU inference is old news

Fahd Mirza profil fotoğrafı
Fahd Mirza8 ay önce

🚨 THIS 100MB TTS MODEL RUNS ON YOUR LAPTOP 🎆 ♠ and it's 6x faster than real-time 🚀 🔹 Pocket-TTS = voice cloning with zero GPU requirements 🔹 CALM architecture skips token generation entirely 🔹 200ms latency + offline + handles infinite text 🔹 How continuous modeling killed expensive APIs 🔥 Watch the full breakdown here:

Ananya Patelik profil fotoğrafı
Ananya Patelik8 ay önce

100M params and no GPU? that’s wizardry. tried it on my old thinkpad and it sings. voice cloning is scarily good. hats off, team!

Amos Gyamfi profil fotoğrafı
Amos Gyamfi8 ay önce

Integrating this with will be awwwesome. I'm eager to try it.

SebastianBoo profil fotoğrafı
SebastianBoo8 ay önce

Can you share the train code? 🥺

jezza kezza profil fotoğrafı
jezza kezza8 ay önce

I use voice clone TTS every day. Fine as an 'AI character' model, well done. Not good enough for things like podcasts, dubbing or back-patching existing recording as it loses some natural character. To be fair, even elevenLabs is borderline for that in professional settings, so no shade for a tiny model ! Maybe OK for promo uses.

Dark profil fotoğrafı
Dark8 ay önce

need to inline it with PDF controls. 95% of the TTS I use is reading PDFs. read highlighted section, page, etc

jim profil fotoğrafı
jim8 ay önce

Got it running in a Docker container and working as my Home Assistant Voice. Much nicer results than Piper and about the same speed!!! Very good

kaiyes profil fotoğrafı
kaiyes8 ay önce

would love to use this in mobile. Any implementation yet ? Also, can it do Dual language at the same time ? Like 2 languages spoken in the same sentence - can it catch that ?

Matthew Man e/acc ⏩ profil fotoğrafı
Matthew Man e/acc ⏩8 ay önce

How does it perform next to VibeVoice? Having a benchmark on your project page would be more legit. There’s so many projects, so when something ships without benchmarks it gets pushed to the “maybe I’ll check this out later” pile. With benchmarks comes priority. FYI

Texas Matt profil fotoğrafı
Texas Matt8 ay önce

post it on r/OpenSourceAI

nle profil fotoğrafı
nle8 ay önce

For normal TTS use cases (not voice cloning), how realistic the sound is comparing to Kokoro? Thanks.

Jim Faster profil fotoğrafı
Jim Faster8 ay önce

Everything that Kyutai does is top tier.

CrossProduct profil fotoğrafı
CrossProduct8 ay önce

Whats time to first audio though ?

IvannP profil fotoğrafı
IvannP8 ay önce

Nice French accent :) « On est Français oui oui baguette »

Frmn profil fotoğrafı
Frmn8 ay önce

vibecoded the web UI using antigravity. really nice, i'm able to generate a bit slower using old thinkpad with i5-1145G7. but with download functions, the audio is clear and good.

Chaitany profil fotoğrafı
Chaitany8 ay önce

This is cool definitely going to try.

Dan profil fotoğrafı
Dan8 ay önce

When you said said MacBook Air, did you mean intel or apple silicon? & if the latter, which M series chip? I’m trying to gauge performance

b90 profil fotoğrafı
b908 ay önce

c’est magnifique!

Hamza profil fotoğrafı
Hamza8 ay önce

@1million_ae

Goblin Town Citizen, CFA --rtrd/acc profil fotoğrafı
Goblin Town Citizen, CFA --rtrd/acc8 ay önce

gpt sovits seems better

ryu profil fotoğrafı
ryu8 ay önce

Have you tried a short context transformer with k=1, i.e. just a skip connection? Makes me think of JiT with the x0 pred v-loss

Zwecharki of Immense AGI fundamentalism profil fotoğrafı
Zwecharki of Immense AGI fundamentalism8 ay önce

@LocallyAIApp wohoo?

Odartei Isaiah profil fotoğrafı
Odartei Isaiah8 ay önce

@grok can this view my ide and help me learn and also learn along with me. How best can it be used

R1K-42 profil fotoğrafı
R1K-428 ay önce

Which languages are planned to be supported? Do you plan to make the code available for fine-tuning?

PanMan profil fotoğrafı
PanMan8 ay önce

@amuldotexe

Idrissi profil fotoğrafı
Idrissi8 ay önce

Congrats 🎆

FoundationModels profil fotoğrafı
FoundationModels8 ay önce

wow just 100M parameters - great training

HappyVolley profil fotoğrafı
HappyVolley8 ay önce

I can't tell you how hard it is to find a good TTS model. Popular ones get silod or paywalled. We know the ones. I've made some real demonic garbled speech before.....onwards kyutai

#INTERNETofAGENTS profil fotoğrafı
#INTERNETofAGENTS8 ay önce

Amazing. 💻

RelakkesYang profil fotoğrafı
RelakkesYang8 ay önce

cool

Giacomo Arienti profil fotoğrafı
Giacomo Arienti8 ay önce

Waiting for a multilingual version, please consider to include Italian 🇮🇹

Benzer Videolar