
kyutai
@kyutai_labs • 27,696 subscribers
Shorts
Videos

If you're looking for a weekend project, how about training your own text-to-speech model from scratch on your own GPU, and then running it on any device's CPU? We just open-sourced the entire Pocket TTS training stack: data pipeline, recipes, and evals. It learns pretty damn fast: ~15k steps: babbling starts turning into words ~50k steps: it reads anything you type (WER under 1%) ~200k steps: the voice stops sounding synthetic On a beefy consumer GPU, that's a week of training. On eight H100s: 10-20 hours. A TTS training run will cost you less than $200 if you rent your hardware, and an order of magnitude less if you just pay for power. Some things we'd love to see people try: - Train it in your own language (a few hundred hours of speech gets you surprisingly far). - Add new features to Pocket TTS (Emotion tags? Make it sing?). - Beat us at our own game: make it faster and smaller. Show us what you build! We'll highlight the best models and new languages for the whole community to enjoy. Pocket TTS has already found many use cases, from reading for people with visual impairments to making NPCs in video games talk, and we're sure there's much more to do with it! Here's an example of a Czech Pocket TTS. Try just asking your favorite agent to find data and apply the method, and you can have your own. Get started:
kyutai103,670 görüntüleme • 9 gün önce

Speech-native models like Moshi sound great and answer fast, but aren’t as smart as text LLMs. In our new paper, MoshiRAG, we show how Moshi can ask for advice from a text LLM or a knowledge base. The tricky part is how to do this in real time without adding latency. 🧵
kyutai53,142 görüntüleme • 4 ay önce

Meet Hibiki, our simultaneous speech-to-speech translation model, currently supporting 🇫🇷➡️🇬🇧. Hibiki produces spoken and text translations of the input speech in real-time, while preserving the speaker’s voice and optimally adapting its pace based on the semantic content of the source speech. Based on objective and human evaluations, Hibiki outperforms previous systems for quality, naturalness and speaker similarity and approaches human interpreters. 🧵
kyutai167,614 görüntüleme • 1 yıl önce

New paper: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models We use RL to post-train speech models (Moshi and PersonaPlex) to talk more like a human: to know when to respond, when to wait, and when to nod along with “yeah”s and “okay”s when listening.
kyutai33,244 görüntüleme • 2 ay önce

Meet MoshiVis🎙️🖼️, the first open-source real-time speech model that can talk about images! It sees, understands, and talks about images — naturally, and out loud. Voice interaction with a compact model endowed with visual understanding opens up new applications, from audio description for the visual impaired to visual access to information. Try it out 👉 Blog post 👉
kyutai47,980 görüntüleme • 1 yıl önce

Thanks to Xavier Niel for stopping by at the #AIActionSummit to try Hibiki. No need to struggle with English anymore 😅
kyutai24,612 görüntüleme • 1 yıl önce

Have you enjoyed talking to 🟢Moshi? Have you dreamt of making your own speech to speech chat experience🧑🔬🤖 ? It's now possible with the moshi-finetune codebase! Plug your own dataset and change the voice, the tone and the personality of Moshi 💚🔌💿. Here's an example after finetuning w/ only 20 hours from the public DailyTalk dataset. 🧵
kyutai20,116 görüntüleme • 1 yıl önce

With Invincible Voice, we help people living with ALS communicate more easily. Encountering Olivier Goy, an entrepreneur who lives with ALS and relentlessly fights to help all patients, made it obvious that our cutting-edge voice AI should help. We turned our Unmute voice-wrapper into a new system that 1/ transcribes interlocutor’s speech in real time, 2/ suggests various relevant responses via a personalised language model, 3/ utters patient's chosen response with their voice (using 10s pre-disease speech recordings). True to our philosophy, we open-source Invincible Voice, so that developers can refine the prototype, port it from French to other languages, adapt it to other conditions (aphasia, neurodegenerative diseases) and turn it into a deployable product. Its modularity also allows it to leverage technologies developed by Gradium that supports Invincible Voice by granting it free access to its multilingual speech models.
kyutai11,106 görüntüleme • 7 ay önce

Even KAVINSKY✨ 🎧🪩 can't break Hibiki! Just like Moshi, Hibiki is robust to extreme background conditions 💥🔊.
kyutai11,866 görüntüleme • 1 yıl önce