正在加载视频...
视频加载失败
Kyutai released their Streaming Text to Speech model, ~2B param model, ultra low latency (220ms), CC-BY-4.0 license 🔥 Trained on 2.5 Million Hours of audio, it can serve up to 32 users w/ less than 350ms latency on a SINGLE L40 🤯 Incredible release by kyutai folks, go check... show more
6 条评论

Vaibhav (VB) Srivastav1 年前
Check out their models here:

Aakash1 年前
"Trained on 2.5 Million Hours of audio, it can serve up to 32 users w/ less than 350ms latency on a SINGLE L40" can we get more of this benchmark

KD1 年前
These are some of the same guys who run a really amazing YT channel about CS btw:

ZAZO1 年前
that’s the best thing happened in 2025 🔥🔥🔥🔥🔥🔥🔥🔥🔥

Bui Dinh Ngoc1 年前
This is game-changing for accessibility tools. I've been waiting for low-latency TTS that doesn't break the bank or require proprietary licenses.

Carlos DP1 年前
SUCH a solid demo lol, S tier
