Loading video...

Video Failed to Load

Go Home

Kyutai Speech-To-Text is now open-source! It’s streaming, supports batched inference, and runs blazingly fast: perfect for interactive applications. Check out the details here:

66,503 views • 1 year ago •via X (Twitter)

9 Comments

kyutai's profile picture
kyutai1 year ago

Today we are releasing two models. The first one is a 2.6B English-only model that beats Whisper Large v3 on benchmarks even though it’s a streaming model that doesn’t process all the audio at once. It can process 400 sequences in parallel on a single H100.

kyutai's profile picture
kyutai1 year ago

The other model is a lightweight English/French 1B model optimized for real-time voice chat apps like It comes with a semantic voice activity detector that predicts if you’re done talking or just pausing mid-sentence. The open-source releases of Kyutai Text-To-Speech and will follow soon!

clem 🤗's profile picture
clem 🤗1 year ago

Magnifique !

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)1 year ago

This is great!! Well cover on @thursdai_pod on an hour

@gerry's profile picture
@gerry1 year ago

That is really good. Well done :)

Dan Western's profile picture
Dan Western1 year ago

Interesting... Great conversation with this ai. Wondering about potential opportunities to embed this functionality into apps...

karai's profile picture
karai1 year ago

It needs mooore languages

ratwell's profile picture
ratwell1 year ago

@dankvr finally

Simon Icard 's profile picture
Simon Icard 1 year ago

👏

Related Videos