Video yükleniyor...
Video Yüklenemedi
The world’s smallest Transformer-based TTS model? We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.
74,884 görüntüleme • 1 ay önce •via X (Twitter)
39 Yorum

What can a 0.1B-class TTS model actually sound like? Listen to the voiceover in this demo video. Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian, Japanese, and Korean Small in scale. Surprisingly capable in speech.

On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves: • English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2 These results are achieved with an approximately 170M-parameter main model.

The model, codec, tokenizer, processor, and inference code are now available: Model: Try it, test the voice cloning capability, and share your feedback.

update: INT8 version only 170MB

Nice work, but I think Pocket-TTS ( is ~100M params.

Thanks. Good catch — Pocket-TTS is indeed smaller at ~100M parameters. Our focus is compact multilingual TTS, zero-shot voice cloning, and efficient end-to-end speech generation.

Tiny footprint, impressive voice quality. Open-source TTS keeps getting more exciting.

170M parameters and sub 2% English WER is wild. Shows multilingual voice cloning can be high quality without a massive model. Size isn’t the only factor for TTS performance

It is possible to run it on mobile like Supertonic ?

Sure, It works.

This is insane 😲

huh wow ~0.2b for this quality?

Kokoro is 82 million. Still excited to check this out!

ElevenLabs executives watching a 170M parameter open-source model do zero-shot voice cloning locally for free: 👁️👄👁️

@Presidentlin

This is the useful shape of an exponential: not a bigger model, a closed loop! Speech that used to need a cluster now fits in 170 million parameters. The compounding is local. The next question is, which other loops have already closed this way and still get talked about as if they need the frontier?

Fair, 170M holding cloud-level voice quality says the ceiling was never really about size, a point we came at from the other side with a floatless architecture rather than a traditional transformer.

@Petucato

Będę pod wrażeniem jak dodacie Polski. Kyutai ma najmniejsze modele. Nikt nie pamięta o języku polskim.

love it!

impressive lets see if my hermes agent can talk to me in squeeky anime voice with Audio8

170M params, cloud-level quality, impressive.

This is incredible. I was just searching for smaller cloning TTS models today and coincidentally found this now.

170M parameters with multilingual TTS and zero shot voice cloning is impressive. The smaller footprint could make local deployment much more practical.

Wow how about phone version?

170m is compelling, but multilingual pronunciation drift matters more than demo voice quality

Useful for students implementing papers too. Before chasing the newest method, build a simple baseline, reproduce the core setting, then add ablations so the improvement is actually explainable.

Awesome 👏

Fi fa fi fi faaa

Did you change the license to non-commercial?

at 0.1B the interesting constraint is prosody. tiny tts models nail phonemes but struggle with long-range intonation, the sentence-level melody. if you kept it natural at that size that's the real achievement, since the payoff is running fully on-device with no latency.

I did a small scale testing, so far the best accuracy for cantonuese is sensevoice and whisper cantonese fine tune, audio8 is close, better than whisper base model

looks and sounds great, isn't kokorro tts 88million parameters though?

The small-footprint part matters a lot for reading tools. Natural TTS is great, but fast enough to use on every article or PDF is where it becomes a habit.

Naming Chinese and English as primary and the rest behind them is more honest than most releases. How wide is the gap between those tiers at 170M? Asking because the open source model I use leaked untranslated lines into dubs, and that was an alignment bug, not size.

The Japanese one sounded cute.

很棒,我要来试试

한국어 쪽에서 이질감을 줄이고 싶다면 성조가 없는(또는 약한) 언어를 구분하고 학습시켜야 할 것 같네요. 한국어 표준말 측면에서는 상당히 이질적으로 들립니다. (거의 모든 원어민이 AI음성이라는 걸 바로 알아챌 정도랄까요?)

How does this fare against miso TTS?
