Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model,...

17,569 Aufrufe • vor 3 Tagen •via X (Twitter)

11 Kommentare

Profilbild von kyutai
kyutaivor 3 Tagen

PocketTTS speaks in continuous latents, one every 80 ms, and a small head turns noise into the next latent in a single forward pass. We used to train with LSD, another one-step flow matching loss, but this required Jacobian-vector products at training time. This is not needed anymore with the drifting recipe.

Profilbild von kyutai
kyutaivor 3 Tagen

The key ingredient we came up with to make it work is a learned kernel temperature. The full write-up with the recipe and all the ablations is here : Code: Original Drifting paper:

Profilbild von Aïna
Aïnavor 3 Tagen

Can I just say this little guy is so cute

Profilbild von LastResort
LastResortvor 3 Tagen

@grok which language it supports?

Profilbild von {stc}
{stc}vor 3 Tagen

nice work playas

Profilbild von Francisco Carlos Erra
Francisco Carlos Erravor 2 Tagen

🙂

Profilbild von Eman Khoubian
Eman Khoubianvor 3 Tagen

less than 1% WER on TTS means a speech recognizer heard the words. it says nothing about whether the cloned voice still sounds right at minute 20, which is where synthetic voices break.

Profilbild von Yazid Janati
Yazid Janativor 3 Tagen

@YentiCollin @samsongvch

Profilbild von Vi VL
Vi VLvor 3 Tagen

Есть поддержка realtime?

Profilbild von mdzor
mdzorvor 3 Tagen

@LocallyAIApp 🙈

Profilbild von semonic
semonicvor 3 Tagen

Awesome i love this project I just wish you guys jad used more and better training data

Ähnliche Videos

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,336 Aufrufe • vor 2 Jahren