正在加载视频...
视频加载失败
MARS5 TTS: Open Source Text to Speech with insane prosodic control! 🔥 > Voice cloning with less than 5 seconds of audio > Two stage Auto-Regressive (750M) + Non-Auto Regressive (450M) model architecture > Used BPE tokenizer to enable control over punctuations, pauses, stops etc. > AR model predicts... show more
10 条评论

Vaibhav (VB) Srivastav2 年前
Check out the model here:

Vaibhav (VB) Srivastav2 年前
GitHub for more deets:

Carlos DP2 年前
Wow, these outputs are incredible. Like, is this the new SOTA? The samples sound better than the 11labs ones, at least, but idk what params were used

Vaibhav (VB) Srivastav2 年前
750M + 450M -> pretty lightweight overall, in the GitHub README they promise more updates coming soon :D

Furkan Gözükara2 年前
5 seconds to clone is always a lie but i can't say for sure without testing i asked them for gradio demo app to be shared

marko.2 年前
Released under GNU AGPL 3.0, a very curious choice for a model but I'll take it 🎉

Marouane Belkouri2 年前
Finnetunning code ?

adivina_soy32 年前
@huggingface Impresionante. Crees que seria posible combinarlo con Hallo?

Thomas Hill2 年前
Nice share 🔥

STEVE blowJOBS2 年前
This is racist ask me why
