正在加载视频...
视频加载失败
We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model,... show more
17,569 次观看 • 3 天前 •via X (Twitter)
11 条评论

PocketTTS speaks in continuous latents, one every 80 ms, and a small head turns noise into the next latent in a single forward pass. We used to train with LSD, another one-step flow matching loss, but this required Jacobian-vector products at training time. This is not needed anymore with the drifting recipe.

The key ingredient we came up with to make it work is a learned kernel temperature. The full write-up with the recipe and all the ablations is here : Code: Original Drifting paper:

Can I just say this little guy is so cute

@grok which language it supports?

nice work playas

🙂

less than 1% WER on TTS means a speech recognizer heard the words. it says nothing about whether the cloned voice still sounds right at minute 20, which is where synthetic voices break.

@YentiCollin @samsongvch

Есть поддержка realtime?

@LocallyAIApp 🙈

Awesome i love this project I just wish you guys jad used more and better training data
