Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Kyutai Speech-To-Text is now open-source! It’s streaming, supports batched inference, and runs blazingly fast: perfect for interactive applications. Check out the details here:

66,503 görüntüleme • 1 yıl önce •via X (Twitter)

9 Yorum

kyutai profil fotoğrafı
kyutai1 yıl önce

Today we are releasing two models. The first one is a 2.6B English-only model that beats Whisper Large v3 on benchmarks even though it’s a streaming model that doesn’t process all the audio at once. It can process 400 sequences in parallel on a single H100.

kyutai profil fotoğrafı
kyutai1 yıl önce

The other model is a lightweight English/French 1B model optimized for real-time voice chat apps like It comes with a semantic voice activity detector that predicts if you’re done talking or just pausing mid-sentence. The open-source releases of Kyutai Text-To-Speech and will follow soon!

clem 🤗 profil fotoğrafı
clem 🤗1 yıl önce

Magnifique !

Alex Volkov (Thursd/AI) profil fotoğrafı
Alex Volkov (Thursd/AI)1 yıl önce

This is great!! Well cover on @thursdai_pod on an hour

@gerry profil fotoğrafı
@gerry1 yıl önce

That is really good. Well done :)

Dan Western profil fotoğrafı
Dan Western1 yıl önce

Interesting... Great conversation with this ai. Wondering about potential opportunities to embed this functionality into apps...

karai profil fotoğrafı
karai1 yıl önce

It needs mooore languages

ratwell profil fotoğrafı
ratwell1 yıl önce

@dankvr finally

Simon Icard  profil fotoğrafı
Simon Icard 1 yıl önce

👏

Benzer Videolar