Загрузка видео...
Не удалось загрузить видео
MARS5 TTS: Open Source Text to Speech with insane prosodic control! 🔥 > Voice cloning with less than 5 seconds of audio > Two stage Auto-Regressive (750M) + Non-Auto Regressive (450M) model architecture > Used BPE tokenizer to enable control over punctuations, pauses, stops etc. > AR model predicts... show more
162,180 просмотров • 2 лет назад •via X (Twitter)
Комментарии: 10

Vaibhav (VB) Srivastav2 лет назад
Check out the model here:

Vaibhav (VB) Srivastav2 лет назад
GitHub for more deets:

Carlos DP2 лет назад
Wow, these outputs are incredible. Like, is this the new SOTA? The samples sound better than the 11labs ones, at least, but idk what params were used

Vaibhav (VB) Srivastav2 лет назад
750M + 450M -> pretty lightweight overall, in the GitHub README they promise more updates coming soon :D

Furkan Gözükara2 лет назад
5 seconds to clone is always a lie but i can't say for sure without testing i asked them for gradio demo app to be shared

marko.2 лет назад
Released under GNU AGPL 3.0, a very curious choice for a model but I'll take it 🎉

Marouane Belkouri2 лет назад
Finnetunning code ?

adivina_soy32 лет назад
@huggingface Impresionante. Crees que seria posible combinarlo con Hallo?

Thomas Hill2 лет назад
Nice share 🔥

STEVE blowJOBS2 лет назад
This is racist ask me why

