Загрузка видео...

Не удалось загрузить видео

На главную

We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model,...

17,569 просмотров • 3 дней назад •via X (Twitter)

Комментарии: 11

Фото профиля kyutai
kyutai3 дней назад

PocketTTS speaks in continuous latents, one every 80 ms, and a small head turns noise into the next latent in a single forward pass. We used to train with LSD, another one-step flow matching loss, but this required Jacobian-vector products at training time. This is not needed anymore with the drifting recipe.

Фото профиля kyutai
kyutai3 дней назад

The key ingredient we came up with to make it work is a learned kernel temperature. The full write-up with the recipe and all the ablations is here : Code: Original Drifting paper:

Фото профиля Aïna
Aïna3 дней назад

Can I just say this little guy is so cute

Фото профиля LastResort
LastResort3 дней назад

@grok which language it supports?

Фото профиля {stc}
{stc}3 дней назад

nice work playas

Фото профиля Francisco Carlos Erra
Francisco Carlos Erra2 дней назад

🙂

Фото профиля Eman Khoubian
Eman Khoubian3 дней назад

less than 1% WER on TTS means a speech recognizer heard the words. it says nothing about whether the cloned voice still sounds right at minute 20, which is where synthetic voices break.

Фото профиля Yazid Janati
Yazid Janati3 дней назад

@YentiCollin @samsongvch

Фото профиля Vi VL
Vi VL3 дней назад

Есть поддержка realtime?

Фото профиля mdzor
mdzor3 дней назад

@LocallyAIApp 🙈

Фото профиля semonic
semonic3 дней назад

Awesome i love this project I just wish you guys jad used more and better training data

Похожие видео

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,336 просмотров • 2 лет назад