Loading video...

Video Failed to Load

Go Home

Introducing Kitten TTS, a SOTA tiny text-to-speech model - Just 15M parameters - Runs without a GPU - Model size less than 25 MB - Multiple high-quality voices - Ultra-fast - even runs on low-end edge devices Github and HF links below

349,016 views • 1 year ago •via X (Twitter)

47 Comments

Divam Gupta's profile picture
Divam Gupta1 year ago

Github:  Huggingface: 

Divam Gupta's profile picture
Divam Gupta1 year ago

Someone deployed it on a browser within few hrs!

Divam Gupta's profile picture
Divam Gupta1 year ago

Also, we would love to collaborate with y’all. If you want to potentially use this model somewhere, shoot me a DM

Paul Bohm's profile picture
Paul Bohm1 year ago

But does it meow?

Divam Gupta's profile picture
Divam Gupta1 year ago

Added it to our todo!

Hamzé 🦀's profile picture
Hamzé 🦀1 year ago

@joshmo_dev We could use on Raspberry pi so?

Divam Gupta's profile picture
Divam Gupta1 year ago

@joshmo_dev Ofc

Hamzé 🦀's profile picture
Hamzé 🦀1 year ago

@joshmo_dev 🤩 wow

Hamzé 🦀's profile picture
Hamzé 🦀1 year ago

@joshmo_dev I would say this is more game changer that gpt-oss

Paweł Bauer's profile picture
Paweł Bauer1 year ago

What languages does it support?

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺's profile picture
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 year ago

I genuinely appreciate your efforts and generosity but is this a case of a 25 MB model that requires gigabytes of python library downloads to run? Any chance in the future this can be ported to something written in C while being as lean and mean as the FFMpeg guys would write?

Divam Gupta's profile picture
Divam Gupta1 year ago

Yes we will do that

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺's profile picture
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 year ago

Awsome! 🙏

Josh Whiton's profile picture
Josh Whiton1 year ago

Wow... so good, so smol. Very nice work!

Divam Gupta's profile picture
Divam Gupta1 year ago

This is an early checkpoint. The model is expected to get better!

Karim C's profile picture
Karim C1 year ago

15M parameters for quality TTS? this is exactly what edge deployment needed been looking for something like this for mario's voice responses in offline scenarios 25MB means it fits on basically any device, no cloud dependency the small model revolution is real

Vishnu Saran's profile picture
Vishnu Saran1 year ago

Great work! Can we also do a custom clone?

Divam Gupta's profile picture
Divam Gupta1 year ago

We will release finetuning for custom voices in the future

Patryk Zoltowski's profile picture
Patryk Zoltowski1 year ago

@ycombinator Couldn't find what languages it supports. Is it only English?

Erik Kaiser's profile picture
Erik Kaiser1 year ago

Dude. I am all over this. This is awesome. Do you have a mobile SDK yet? I have an ESP32 application for this right now today.

Divam Gupta's profile picture
Divam Gupta1 year ago

No sdk yet. But it’s just onnx. Maybe we can work together to get it on ESP32

Nasim Uddin's profile picture
Nasim Uddin1 year ago

Wow. Nice. Will there be other SDK for running in mobile?

Divam Gupta's profile picture
Divam Gupta1 year ago

Definitely! Coming soon

Andromedus's profile picture
Andromedus1 year ago

Are you guys by any chance working on a tiny speech-to-text model?

Divam Gupta's profile picture
Divam Gupta1 year ago

We are!

Pendar's profile picture
Pendar1 year ago

This is awesome, I can now add offline voice capability to my app

Divam Gupta's profile picture
Divam Gupta1 year ago

Would love for you to add it! Dm me

Ari Kouts's profile picture
Ari Kouts1 year ago

So it's raspberry friendly?

Gordon Shumway's profile picture
Gordon Shumway1 year ago

Holy moly. This is the most important announcement today if voice output is generally this good for everything and it wasn’t cherry picked! And yes, I’m aware of GPT OSS and Opus 4.1 releases, but they will be replaced with something better in a week or so and this tiny tts model with audio output this good, is way more important if you ask me!

ElonBald's profile picture
ElonBald1 year ago

Release the training code!!

Venkata Pingali's profile picture
Venkata Pingali1 year ago

@divamgupta Great work. Was trying it to make it work on larger segments of text. It is failing even at 1000 characters. Shorter ones work. Same issue as

Satish G's profile picture
Satish G1 year ago

Wow! Gonna try it out today and let you know my feedback.

mt_shammah's profile picture
mt_shammah1 year ago

@Hamzeml Awesome work 25mb is crazy. Needed something this small a few months ago. Will try this out soon

Privacy AI - on-device AI agent & chatbot's profile picture
Privacy AI - on-device AI agent & chatbot1 year ago

It’s incredible to achieve this with only 15M parameters! We’re integrating Kokoros’ 82M model into our application now, and your model is on par with it. Congratulations! 🎉

David Alejandro's profile picture
David Alejandro1 year ago

Awesome! What’s the difference between this and just using a plain ol’ TTS system? Voices seem similar. Genuinely asking.

Nikita 🤙's profile picture
Nikita 🤙1 year ago

Great work! Will try in my TTS engine with livekit

Stefano Rivera's profile picture
Stefano Rivera1 year ago

@NeoFablesVR check this out.

Yeab - engg / core was in Palo Alto🌴's profile picture
Yeab - engg / core was in Palo Alto🌴1 year ago

smaller

Prashant's profile picture
Prashant1 year ago

Looking so good.. 👏

Chris Matthieu's profile picture
Chris Matthieu1 year ago

Sounds great!

Devin AI's profile picture
Devin AI1 year ago

Innovative work! Edge AI models like Kitten TTS make tech accessible. Exciting to see AI running without heavy hardware requirements.

Somesh's profile picture
Somesh1 year ago

great stuff!

Nathaniel™'s profile picture
Nathaniel™1 year ago

I need an Android PDF reader to have this as an add-on

Dimitri Gilbert's profile picture
Dimitri Gilbert1 year ago

you sir, got my interest ! :) been a long day but it is bm for tomorrow !

@StoryOfAIGuess's profile picture
@StoryOfAIGuess1 year ago

Very cool man great success

Fareesh Vijayarangam's profile picture
Fareesh Vijayarangam1 year ago

any plans for cloning? accents?

inlanger 🍻 🇺🇦's profile picture
inlanger 🍻 🇺🇦1 year ago

Should it work in a browser?

Related Videos