Video yükleniyor...
Video Yüklenemedi
Today we're releasing ZONOS2, our next-generation real-time TTS model with high-fidelity voice cloning. ZONOS2 is the most expressive open-source TTS model, released under Apache 2.0 and available on Zyphra Cloud on AMD. 🧵
334,634 görüntüleme • 3 ay önce •via X (Twitter)
25 Yorum

Real-time TTS has always forced a tradeoff between quality and speed. We achieve both with ZONOS2, the first sparse MoE TTS model released open-source, with 8B total params, 900M active. ZONOS2 is fast, inference efficient, and super expressive.

ZONOS2 excels at voice cloning, making it the most natural-sounding open-source TTS model out there. It captures far more of what makes a voice distinctive, so clones sound convincing across a wide range of speakers. Voice cloning is zero-shot, needing no fine-tuning.

ZONOS2 predicts Descript Audio Codec (DAC) tokens for studio-quality 44.1 kHz audio. DAC tokens maximize quality but are harder to model than low-fi autoencoders. We close that gap with model + data scale, so fidelity doesn't cost stability.

For the text, we do not use a phonemizer, instead ZONOS2 reads raw UTF-8 bytes. This gives us: → broader coverage, especially lower-resource languages → big gains on Chinese, Korean, Japanese → native code-switching mid-sentence

Training data scaled from ~200K hours to 6M+ hours (~707 years of audio). Staged data filtering ramps transcript-agreement strictness across pretraining → midtraining → annealing. This leads to fewer hallucinations, mispronunciations, and repetitions.

We're also releasing ZTTS1-Eval, a new TTS benchmark. Existing evals lean on outdated ASR and read speech. ZTTS1-Eval spans clean + in-the-wild sets across up to 17 languages, modern judges (Qwen3-ASR, ReDimNet, MSR-UTMOS), and prosody metrics.

ZONOS2 is open-weights under Apache 2.0, and free on Zyphra Cloud for a limited time. Try it on Zyphra Cloud: Blog: Weights: Inference code: Eval code:

is an open superintelligence research and product company based in San Francisco, CA on a mission to build human-aligned AI that helps individuals and organizations reach their fullest potential. Apply to join us!

@AMD >available on Zyphra Cloud on @AMD. >Platform Support: Linux only (x86 64). Requires NVIDIA GPU with CUDA toolkit matching your driver version (nvidia- smi to check)

@AMD "controllable, safe and aligned" -- what, so it can't use bad words?

@AMD Is the current audio playground on Zyphra Cloud using Zonos2?

@AMD Idk what benchmarks these are but the voice capabilities definitely don’t sound Gemini Flash level man cmon.

@AMD sounds like shit

@AMD damn that elon and sam voice is so on point, great work team

@AMD ZONOS2 pushes open-source TTS into high fidelity, expressive territory

@AMD what's the cost to run it on zyphra cloud?

@AMD Congrats! 🔥

@AMD speaker drift over long sessions kills voice clone trust faster than headline fidelity gains

@AMD ZONOS2 is a huge leap for open-source TTS—real-time, highly expressive, and pushing voice AI innovation to the next level!

@AMD congrats!!

@AMD Congrats team!

@AMD Big release 🔥 real-time + expressive voice cloning is a killer combo. Love that it’s open-source too, this is gonna unlock a lot

@AMD Open-source real-time TTS is especially useful when latency, expressiveness, and deployment flexibility improve together.

@AMD odd, I checked the repo and im sure I saw that this could only be run on nvidia so far. I’m also on an AMD (W7900), how can I run inference on this?

@AMD How about ZONOS2 on local AMD?
