Loading video...
Video Failed to Load
MARS5 TTS: Open Source Text to Speech with insane prosodic control! 🔥 > Voice cloning with less than 5 seconds of audio > Two stage Auto-Regressive (750M) + Non-Auto Regressive (450M) model architecture > Used BPE tokenizer to enable control over punctuations, pauses, stops etc. > AR model predicts... show more
162,281 views • 2 years ago •via X (Twitter)
10 Comments

Vaibhav (VB) Srivastav2 years ago
Check out the model here:

Vaibhav (VB) Srivastav2 years ago
GitHub for more deets:

Carlos DP2 years ago
Wow, these outputs are incredible. Like, is this the new SOTA? The samples sound better than the 11labs ones, at least, but idk what params were used

Vaibhav (VB) Srivastav2 years ago
750M + 450M -> pretty lightweight overall, in the GitHub README they promise more updates coming soon :D

Furkan Gözükara2 years ago
5 seconds to clone is always a lie but i can't say for sure without testing i asked them for gradio demo app to be shared

marko.2 years ago
Released under GNU AGPL 3.0, a very curious choice for a model but I'll take it 🎉

Marouane Belkouri2 years ago
Finnetunning code ?

adivina_soy32 years ago
@huggingface Impresionante. Crees que seria posible combinarlo con Hallo?

Thomas Hill2 years ago
Nice share 🔥

STEVE blowJOBS2 years ago
This is racist ask me why
