Загрузка видео...

Не удалось загрузить видео

На главную

We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required. Open-source, lightweight, and incredibly fast. 🧵👇

238,618 просмотров • 8 месяцев назад •via X (Twitter)

Комментарии: 40

Фото профиля kyutai
kyutai8 месяцев назад

The gap in TTS today: ✘ Huge LLM models (1B+ params) need a GPU. ✘ Small models like Kokoro (82M) are fast but lack flexible voice cloning. Pocket TTS bridges the gap. It runs faster than real-time on an average laptop’s CPU while keeping the power of the giants.

Фото профиля kyutai
kyutai8 месяцев назад

🎙️ True voice cloning: Pocket TTS only needs 5 seconds of audio to capture: ✔ Individual voice "color" ✔ the emotion and accent ✔ Acoustic conditions (reverb, mic quality) Use our voice library or clone any voice from a tiny sample.

Фото профиля kyutai
kyutai8 месяцев назад

📊 The stats don't lie. Despite its size (100M params), Pocket TTS beats F5-TTS and DSM in Word Error Rate (1.84) and Audio Quality ELO. It’s the only model in its class that offers voice cloning *and* runs comfortably on a CPU.

Фото профиля kyutai
kyutai8 месяцев назад

🏗️ How? We moved away from discrete tokens. Pocket TTS is based on the Continuous Audio Language Models (CALM) framework and predicts sequences of continuous latents directly, using a 1-step sampling method (Lagrangian Self-Distillation). CALM paper:

Фото профиля kyutai
kyutai8 месяцев назад

🔓 Open source for everyone. Trained on 88k hours of public English data to ensure reproducibility. Check out the code and the technical breakdown here:

Фото профиля Raj Breno
Raj Breno8 месяцев назад

Is it multilingual?

Фото профиля kyutai
kyutai8 месяцев назад

English-only for now.

Фото профиля Steve grins
Steve grins8 месяцев назад

ran a few samples on my m2 air. startup is instant and voices are scarily close to the original. wild to see this quality without a gpu or cloud. nice work

Фото профиля Seba
Seba8 месяцев назад

CPU inference is old news

Фото профиля Fahd Mirza
Fahd Mirza8 месяцев назад

🚨 THIS 100MB TTS MODEL RUNS ON YOUR LAPTOP 🎆 ♠ and it's 6x faster than real-time 🚀 🔹 Pocket-TTS = voice cloning with zero GPU requirements 🔹 CALM architecture skips token generation entirely 🔹 200ms latency + offline + handles infinite text 🔹 How continuous modeling killed expensive APIs 🔥 Watch the full breakdown here:

Фото профиля Ananya Patelik
Ananya Patelik8 месяцев назад

100M params and no GPU? that’s wizardry. tried it on my old thinkpad and it sings. voice cloning is scarily good. hats off, team!

Фото профиля Amos Gyamfi
Amos Gyamfi8 месяцев назад

Integrating this with will be awwwesome. I'm eager to try it.

Фото профиля SebastianBoo
SebastianBoo8 месяцев назад

Can you share the train code? 🥺

Фото профиля jezza kezza
jezza kezza8 месяцев назад

I use voice clone TTS every day. Fine as an 'AI character' model, well done. Not good enough for things like podcasts, dubbing or back-patching existing recording as it loses some natural character. To be fair, even elevenLabs is borderline for that in professional settings, so no shade for a tiny model ! Maybe OK for promo uses.

Фото профиля Dark
Dark8 месяцев назад

need to inline it with PDF controls. 95% of the TTS I use is reading PDFs. read highlighted section, page, etc

Фото профиля jim
jim8 месяцев назад

Got it running in a Docker container and working as my Home Assistant Voice. Much nicer results than Piper and about the same speed!!! Very good

Фото профиля kaiyes
kaiyes8 месяцев назад

would love to use this in mobile. Any implementation yet ? Also, can it do Dual language at the same time ? Like 2 languages spoken in the same sentence - can it catch that ?

Фото профиля Matthew Man e/acc ⏩
Matthew Man e/acc ⏩8 месяцев назад

How does it perform next to VibeVoice? Having a benchmark on your project page would be more legit. There’s so many projects, so when something ships without benchmarks it gets pushed to the “maybe I’ll check this out later” pile. With benchmarks comes priority. FYI

Фото профиля Texas Matt
Texas Matt8 месяцев назад

post it on r/OpenSourceAI

Фото профиля nle
nle8 месяцев назад

For normal TTS use cases (not voice cloning), how realistic the sound is comparing to Kokoro? Thanks.

Фото профиля Jim Faster
Jim Faster8 месяцев назад

Everything that Kyutai does is top tier.

Фото профиля CrossProduct
CrossProduct8 месяцев назад

Whats time to first audio though ?

Фото профиля IvannP
IvannP8 месяцев назад

Nice French accent :) « On est Français oui oui baguette »

Фото профиля Frmn
Frmn8 месяцев назад

vibecoded the web UI using antigravity. really nice, i'm able to generate a bit slower using old thinkpad with i5-1145G7. but with download functions, the audio is clear and good.

Фото профиля Chaitany
Chaitany8 месяцев назад

This is cool definitely going to try.

Фото профиля Dan
Dan8 месяцев назад

When you said said MacBook Air, did you mean intel or apple silicon? & if the latter, which M series chip? I’m trying to gauge performance

Фото профиля b90
b908 месяцев назад

c’est magnifique!

Фото профиля Hamza
Hamza8 месяцев назад

@1million_ae

Фото профиля Goblin Town Citizen, CFA --rtrd/acc
Goblin Town Citizen, CFA --rtrd/acc8 месяцев назад

gpt sovits seems better

Фото профиля ryu
ryu8 месяцев назад

Have you tried a short context transformer with k=1, i.e. just a skip connection? Makes me think of JiT with the x0 pred v-loss

Фото профиля Zwecharki of Immense AGI fundamentalism
Zwecharki of Immense AGI fundamentalism8 месяцев назад

@LocallyAIApp wohoo?

Фото профиля Odartei Isaiah
Odartei Isaiah8 месяцев назад

@grok can this view my ide and help me learn and also learn along with me. How best can it be used

Фото профиля R1K-42
R1K-428 месяцев назад

Which languages are planned to be supported? Do you plan to make the code available for fine-tuning?

Фото профиля PanMan
PanMan8 месяцев назад

@amuldotexe

Фото профиля Idrissi
Idrissi8 месяцев назад

Congrats 🎆

Фото профиля FoundationModels
FoundationModels8 месяцев назад

wow just 100M parameters - great training

Фото профиля HappyVolley
HappyVolley8 месяцев назад

I can't tell you how hard it is to find a good TTS model. Popular ones get silod or paywalled. We know the ones. I've made some real demonic garbled speech before.....onwards kyutai

Фото профиля #INTERNETofAGENTS
#INTERNETofAGENTS8 месяцев назад

Amazing. 💻

Фото профиля RelakkesYang
RelakkesYang8 месяцев назад

cool

Фото профиля Giacomo Arienti
Giacomo Arienti8 месяцев назад

Waiting for a multilingual version, please consider to include Italian 🇮🇹

Похожие видео