Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model,...

17,569 görüntüleme • 3 gün önce •via X (Twitter)

11 Yorum

kyutai profil fotoğrafı
kyutai3 gün önce

PocketTTS speaks in continuous latents, one every 80 ms, and a small head turns noise into the next latent in a single forward pass. We used to train with LSD, another one-step flow matching loss, but this required Jacobian-vector products at training time. This is not needed anymore with the drifting recipe.

kyutai profil fotoğrafı
kyutai3 gün önce

The key ingredient we came up with to make it work is a learned kernel temperature. The full write-up with the recipe and all the ablations is here : Code: Original Drifting paper:

Aïna profil fotoğrafı
Aïna3 gün önce

Can I just say this little guy is so cute

LastResort profil fotoğrafı
LastResort3 gün önce

@grok which language it supports?

{stc} profil fotoğrafı
{stc}3 gün önce

nice work playas

Francisco Carlos Erra profil fotoğrafı
Francisco Carlos Erra2 gün önce

🙂

Eman Khoubian profil fotoğrafı
Eman Khoubian3 gün önce

less than 1% WER on TTS means a speech recognizer heard the words. it says nothing about whether the cloned voice still sounds right at minute 20, which is where synthetic voices break.

Yazid Janati profil fotoğrafı
Yazid Janati3 gün önce

@YentiCollin @samsongvch

Vi VL profil fotoğrafı
Vi VL3 gün önce

Есть поддержка realtime?

mdzor profil fotoğrafı
mdzor3 gün önce

@LocallyAIApp 🙈

semonic profil fotoğrafı
semonic3 gün önce

Awesome i love this project I just wish you guys jad used more and better training data

Benzer Videolar

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,336 görüntüleme • 2 yıl önce