ๆญฃๅจๅ ่ฝฝ่ง้ข...
่ง้ขๅ ่ฝฝๅคฑ่ดฅ
Announcing ๐๐จ๐ข๐๐๐๐ซ๐๐๐ญ๐ช SotA for both speech editing and zero-shot text-to-speech, Outperforming VALL-E, XTTS-v2, etc. VoiceCraft works on in-the-wild data such as movies, random videos and podcasts We fully open source it at
160,424 ๆฌก่ง็ โข 2 ๅนดๅ โขvia X (Twitter)
10 ๆก่ฏ่ฎบ

๐๐จ๐ข๐๐๐๐ซ๐๐๐ญ works well on recordings with diverse accents, emotions, styles, content, background noise, recording conditions. Demo: Paper: code, model, data:

@rdesh26 Wow! love itt! I'd love to add this on the TTS Arena: Let's chat in DMs?

@rdesh26 Yea letโs do that!

Thank you for releasing! Any possibility of switching to an open source license?

Thanks! Have been discussing the licensing issue, might change it in the coming days

I tried out. Insane!!! Amazing work @PuyuanPeng Can the license be commercially available? Would love to incorporate this into our product.

Cool practical solutions done in sampling: discarding longest/shortest, special cases for silence tokens, rep penalties and silence tokens interacting. Might be able to even penalize the scratching sound if it's consistent enough (or maybe CFG negative prompt?).

Wish it had a more permissive license. Is it possible to change it to be used commercially?

have been discussing the licensing issue, might change it in the coming days

Great work! Would love to connect with you
