正在加载视频...
视频加载失败
1/ First, how do you benchmark an AI model? You give different models the same test cases, and score them. You can do the same with voice models using Word Error Rate (WER). You generate the audio, transcribe it with a speech recognition (ASR) model, and compare that transcript... show more
10,184 次观看 • 23 天前 •via X (Twitter)
7 条评论

We're rigorous about model evals at Cartesia because there's no shortcut. Quality is multidimensional, and no single number captures it. Here’s why benchmarking TTS is extremely difficult ⬇

2/ WER has another problem. The score depends on the ASR model you use to transcribe the audio. If the ASR struggles with an accent, language or domain-specific word, the TTS model gets penalized for the ASR's mistake. The inverse can happen too. A strong ASR can infer unclear speech and make a bad generation look better than it sounds.

3/ Which means we should also measure whether the audio actually sounds good. One way to do that at scale is to get another model to listen to the audio and predict its quality. But these automated quality scores have blind (deaf) spots too. In one study, researchers added increasingly wrong Japanese pitch accents to generated speech. Human ratings dropped from 4.00 to 2.16. But the automated scores didn’t move, or even went up in some cases.

4/ Which is why human evals are still the gold standard. Artificial Analysis, for example, gets people to listen to two samples and choose which one they prefer. Those preferences are then used to rank the models. But human scores aren't completely stable. Results can change based on who's listening, the instructions they're given, headphones vs speakers, volume, or even how tired they are.

5/ Then there's what you actually test. Most TTS evals use short clips because they're easier to generate, listen to and score. But a model can sound great for a few seconds and start breaking over a longer conversation. The speaker can drift, or the accent and pace can change. And when speech is streamed in real time, you get completely different failures like awkward joins, repeated words or cut-off sounds.

6/ A model can also perform really well on generic speech and fall apart on medical terms, legal text, chemistry, gaming or customer support. Each domain has different expectations around pronunciation, pacing and what’s correct. And gets harder across languages, where what counts as an “error” can change by language and locale. So being good on one benchmark doesn't necessarily mean being good on your use case.

7/ So don't just read the leaderboard. Listen for yourself. Test the model for the language, context and product you're actually building. That being said, we're still #1 though :)
