Video wird geladen...
Video konnte nicht geladen werden
Introducing Kitten TTS, a SOTA tiny text-to-speech model - Just 15M parameters - Runs without a GPU - Model size less than 25 MB - Multiple high-quality voices - Ultra-fast - even runs on low-end edge devices Github and HF links below
349,016 Aufrufe • vor 1 Jahr •via X (Twitter)
47 Kommentare

Github: Huggingface:

Someone deployed it on a browser within few hrs!

Also, we would love to collaborate with y’all. If you want to potentially use this model somewhere, shoot me a DM

But does it meow?

Added it to our todo!

@joshmo_dev We could use on Raspberry pi so?

@joshmo_dev Ofc

@joshmo_dev 🤩 wow

@joshmo_dev I would say this is more game changer that gpt-oss

What languages does it support?

I genuinely appreciate your efforts and generosity but is this a case of a 25 MB model that requires gigabytes of python library downloads to run? Any chance in the future this can be ported to something written in C while being as lean and mean as the FFMpeg guys would write?

Yes we will do that

Awsome! 🙏

Wow... so good, so smol. Very nice work!

This is an early checkpoint. The model is expected to get better!

15M parameters for quality TTS? this is exactly what edge deployment needed been looking for something like this for mario's voice responses in offline scenarios 25MB means it fits on basically any device, no cloud dependency the small model revolution is real

Great work! Can we also do a custom clone?

We will release finetuning for custom voices in the future

@ycombinator Couldn't find what languages it supports. Is it only English?

Dude. I am all over this. This is awesome. Do you have a mobile SDK yet? I have an ESP32 application for this right now today.

No sdk yet. But it’s just onnx. Maybe we can work together to get it on ESP32

Wow. Nice. Will there be other SDK for running in mobile?

Definitely! Coming soon

Are you guys by any chance working on a tiny speech-to-text model?

We are!

This is awesome, I can now add offline voice capability to my app

Would love for you to add it! Dm me

So it's raspberry friendly?

Holy moly. This is the most important announcement today if voice output is generally this good for everything and it wasn’t cherry picked! And yes, I’m aware of GPT OSS and Opus 4.1 releases, but they will be replaced with something better in a week or so and this tiny tts model with audio output this good, is way more important if you ask me!

Release the training code!!

@divamgupta Great work. Was trying it to make it work on larger segments of text. It is failing even at 1000 characters. Shorter ones work. Same issue as

Wow! Gonna try it out today and let you know my feedback.

@Hamzeml Awesome work 25mb is crazy. Needed something this small a few months ago. Will try this out soon

It’s incredible to achieve this with only 15M parameters! We’re integrating Kokoros’ 82M model into our application now, and your model is on par with it. Congratulations! 🎉

Awesome! What’s the difference between this and just using a plain ol’ TTS system? Voices seem similar. Genuinely asking.

Great work! Will try in my TTS engine with livekit

@NeoFablesVR check this out.

smaller

Looking so good.. 👏

Sounds great!

Innovative work! Edge AI models like Kitten TTS make tech accessible. Exciting to see AI running without heavy hardware requirements.

great stuff!

I need an Android PDF reader to have this as an add-on

you sir, got my interest ! :) been a long day but it is bm for tomorrow !

Very cool man great success

any plans for cloning? accents?

Should it work in a browser?
