正在加载视频...

视频加载失败

Introducing Kitten TTS, a SOTA tiny text-to-speech model - Just 15M parameters - Runs without a GPU - Model size less than 25 MB - Multiple high-quality voices - Ultra-fast - even runs on low-end edge devices Github and HF links below

349,016 次观看 • 1 年前 •via X (Twitter)

47 条评论

Divam Gupta 的头像
Divam Gupta1 年前

Github:  Huggingface: 

Divam Gupta 的头像
Divam Gupta1 年前

Someone deployed it on a browser within few hrs!

Divam Gupta 的头像
Divam Gupta1 年前

Also, we would love to collaborate with y’all. If you want to potentially use this model somewhere, shoot me a DM

Paul Bohm 的头像
Paul Bohm1 年前

But does it meow?

Divam Gupta 的头像
Divam Gupta1 年前

Added it to our todo!

Hamzé 🦀 的头像
Hamzé 🦀1 年前

@joshmo_dev We could use on Raspberry pi so?

Divam Gupta 的头像
Divam Gupta1 年前

@joshmo_dev Ofc

Hamzé 🦀 的头像
Hamzé 🦀1 年前

@joshmo_dev 🤩 wow

Hamzé 🦀 的头像
Hamzé 🦀1 年前

@joshmo_dev I would say this is more game changer that gpt-oss

Paweł Bauer 的头像
Paweł Bauer1 年前

What languages does it support?

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺 的头像
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 年前

I genuinely appreciate your efforts and generosity but is this a case of a 25 MB model that requires gigabytes of python library downloads to run? Any chance in the future this can be ported to something written in C while being as lean and mean as the FFMpeg guys would write?

Divam Gupta 的头像
Divam Gupta1 年前

Yes we will do that

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺 的头像
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 年前

Awsome! 🙏

Josh Whiton 的头像
Josh Whiton1 年前

Wow... so good, so smol. Very nice work!

Divam Gupta 的头像
Divam Gupta1 年前

This is an early checkpoint. The model is expected to get better!

Karim C 的头像
Karim C1 年前

15M parameters for quality TTS? this is exactly what edge deployment needed been looking for something like this for mario's voice responses in offline scenarios 25MB means it fits on basically any device, no cloud dependency the small model revolution is real

Vishnu Saran 的头像
Vishnu Saran1 年前

Great work! Can we also do a custom clone?

Divam Gupta 的头像
Divam Gupta1 年前

We will release finetuning for custom voices in the future

Patryk Zoltowski 的头像
Patryk Zoltowski1 年前

@ycombinator Couldn't find what languages it supports. Is it only English?

Erik Kaiser 的头像
Erik Kaiser1 年前

Dude. I am all over this. This is awesome. Do you have a mobile SDK yet? I have an ESP32 application for this right now today.

Divam Gupta 的头像
Divam Gupta1 年前

No sdk yet. But it’s just onnx. Maybe we can work together to get it on ESP32

Nasim Uddin 的头像
Nasim Uddin1 年前

Wow. Nice. Will there be other SDK for running in mobile?

Divam Gupta 的头像
Divam Gupta1 年前

Definitely! Coming soon

Andromedus 的头像
Andromedus1 年前

Are you guys by any chance working on a tiny speech-to-text model?

Divam Gupta 的头像
Divam Gupta1 年前

We are!

Pendar 的头像
Pendar1 年前

This is awesome, I can now add offline voice capability to my app

Divam Gupta 的头像
Divam Gupta1 年前

Would love for you to add it! Dm me

Ari Kouts 的头像
Ari Kouts1 年前

So it's raspberry friendly?

Gordon Shumway 的头像
Gordon Shumway1 年前

Holy moly. This is the most important announcement today if voice output is generally this good for everything and it wasn’t cherry picked! And yes, I’m aware of GPT OSS and Opus 4.1 releases, but they will be replaced with something better in a week or so and this tiny tts model with audio output this good, is way more important if you ask me!

ElonBald 的头像
ElonBald1 年前

Release the training code!!

Venkata Pingali 的头像
Venkata Pingali1 年前

@divamgupta Great work. Was trying it to make it work on larger segments of text. It is failing even at 1000 characters. Shorter ones work. Same issue as

Satish G 的头像
Satish G1 年前

Wow! Gonna try it out today and let you know my feedback.

mt_shammah 的头像
mt_shammah1 年前

@Hamzeml Awesome work 25mb is crazy. Needed something this small a few months ago. Will try this out soon

Privacy AI - on-device AI agent & chatbot 的头像
Privacy AI - on-device AI agent & chatbot1 年前

It’s incredible to achieve this with only 15M parameters! We’re integrating Kokoros’ 82M model into our application now, and your model is on par with it. Congratulations! 🎉

David Alejandro 的头像
David Alejandro1 年前

Awesome! What’s the difference between this and just using a plain ol’ TTS system? Voices seem similar. Genuinely asking.

Nikita 🤙 的头像
Nikita 🤙1 年前

Great work! Will try in my TTS engine with livekit

Stefano Rivera 的头像
Stefano Rivera1 年前

@NeoFablesVR check this out.

Yeab - engg / core was in Palo Alto🌴 的头像
Yeab - engg / core was in Palo Alto🌴1 年前

smaller

Prashant 的头像
Prashant1 年前

Looking so good.. 👏

Chris Matthieu 的头像
Chris Matthieu1 年前

Sounds great!

Devin AI 的头像
Devin AI1 年前

Innovative work! Edge AI models like Kitten TTS make tech accessible. Exciting to see AI running without heavy hardware requirements.

Somesh 的头像
Somesh1 年前

great stuff!

Nathaniel™ 的头像
Nathaniel™1 年前

I need an Android PDF reader to have this as an add-on

Dimitri Gilbert 的头像
Dimitri Gilbert1 年前

you sir, got my interest ! :) been a long day but it is bm for tomorrow !

@StoryOfAIGuess 的头像
@StoryOfAIGuess1 年前

Very cool man great success

Fareesh Vijayarangam 的头像
Fareesh Vijayarangam1 年前

any plans for cloning? accents?

inlanger 🍻 🇺🇦 的头像
inlanger 🍻 🇺🇦1 年前

Should it work in a browser?

相关视频