Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

The world’s smallest Transformer-based TTS model? We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.

74,884 Aufrufe • vor 1 Monat •via X (Twitter)

39 Kommentare

Profilbild von Samuel Zeng
Samuel Zengvor 1 Monat

What can a 0.1B-class TTS model actually sound like? Listen to the voiceover in this demo video. Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian, Japanese, and Korean Small in scale. Surprisingly capable in speech.

Profilbild von Samuel Zeng
Samuel Zengvor 1 Monat

On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves: • English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2 These results are achieved with an approximately 170M-parameter main model.

Profilbild von Samuel Zeng
Samuel Zengvor 1 Monat

The model, codec, tokenizer, processor, and inference code are now available: Model: Try it, test the voice cloning capability, and share your feedback.

Profilbild von Samuel Zeng
Samuel Zengvor 26 Tagen

update: INT8 version only 170MB

Profilbild von Desh Raj
Desh Rajvor 1 Monat

Nice work, but I think Pocket-TTS ( is ~100M params.

Profilbild von Samuel Zeng
Samuel Zengvor 1 Monat

Thanks. Good catch — Pocket-TTS is indeed smaller at ~100M parameters. Our focus is compact multilingual TTS, zero-shot voice cloning, and efficient end-to-end speech generation.

Profilbild von RAZA | AI EXPLORER
RAZA | AI EXPLORERvor 1 Monat

Tiny footprint, impressive voice quality. Open-source TTS keeps getting more exciting.

Profilbild von Moez Zhioua
Moez Zhiouavor 1 Monat

170M parameters and sub 2% English WER is wild. Shows multilingual voice cloning can be high quality without a massive model. Size isn’t the only factor for TTS performance

Profilbild von Romain
Romainvor 1 Monat

It is possible to run it on mobile like Supertonic ?

Profilbild von Samuel Zeng
Samuel Zengvor 1 Monat

Sure, It works.

Profilbild von Habtamu Asefa
Habtamu Asefavor 1 Monat

This is insane 😲

Profilbild von Ha Hoang
Ha Hoangvor 1 Monat

huh wow ~0.2b for this quality?

Profilbild von Math3matica
Math3maticavor 1 Monat

Kokoro is 82 million. Still excited to check this out!

Profilbild von Internet Labs
Internet Labsvor 1 Monat

ElevenLabs executives watching a 170M parameter open-source model do zero-shot voice cloning locally for free: 👁️👄👁️

Profilbild von luis
luisvor 1 Monat

@Presidentlin

Profilbild von David Orban
David Orbanvor 1 Monat

This is the useful shape of an exponential: not a bigger model, a closed loop! Speech that used to need a cluster now fits in 170 million parameters. The compounding is local. The next question is, which other loops have already closed this way and still get talked about as if they need the frontier?

Profilbild von Forlais
Forlaisvor 1 Monat

Fair, 170M holding cloud-level voice quality says the ceiling was never really about size, a point we came at from the other side with a floatless architecture rather than a traditional transformer.

Profilbild von Piyush
Piyushvor 1 Monat

@Petucato

Profilbild von 🦀ꃳꁲꌅꋖ🦀
🦀ꃳꁲꌅꋖ🦀vor 1 Monat

Będę pod wrażeniem jak dodacie Polski. Kyutai ma najmniejsze modele. Nikt nie pamięta o języku polskim.

Profilbild von Mahimai Raja J ‎
Mahimai Raja J ‎vor 1 Monat

love it!

Profilbild von kfant
kfantvor 1 Monat

impressive lets see if my hermes agent can talk to me in squeeky anime voice with Audio8

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

170M params, cloud-level quality, impressive.

Profilbild von Human Being
Human Beingvor 1 Monat

This is incredible. I was just searching for smaller cloning TTS models today and coincidentally found this now.

Profilbild von Eva Morgan
Eva Morganvor 1 Monat

170M parameters with multilingual TTS and zero shot voice cloning is impressive. The smaller footprint could make local deployment much more practical.

Profilbild von Johnny Depp
Johnny Deppvor 1 Monat

Wow how about phone version?

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 1 Monat

170m is compelling, but multilingual pronunciation drift matters more than demo voice quality

Profilbild von cordivai | Machine Learning & AI
cordivai | Machine Learning & AIvor 1 Monat

Useful for students implementing papers too. Before chasing the newest method, build a simple baseline, reproduce the core setting, then add ablations so the improvement is actually explainable.

Profilbild von Temel Günaydın
Temel Günaydınvor 1 Monat

Awesome 👏

Profilbild von Rithul
Rithulvor 1 Monat

Fi fa fi fi faaa

Profilbild von Mattias Johansson
Mattias Johanssonvor 1 Monat

Did you change the license to non-commercial?

Profilbild von rusa
rusavor 1 Monat

at 0.1B the interesting constraint is prosody. tiny tts models nail phonemes but struggle with long-range intonation, the sentence-level melody. if you kept it natural at that size that's the real achievement, since the payoff is running fully on-device with no latency.

Profilbild von Beeno Tung
Beeno Tungvor 1 Monat

I did a small scale testing, so far the best accuracy for cantonuese is sensevoice and whisper cantonese fine tune, audio8 is close, better than whisper base model

Profilbild von Michael
Michaelvor 28 Tagen

looks and sounds great, isn't kokorro tts 88million parameters though?

Profilbild von ThuNoJoy | SayTXT for ADHD Reading
ThuNoJoy | SayTXT for ADHD Readingvor 1 Monat

The small-footprint part matters a lot for reading tools. Natural TTS is great, but fast enough to use on every article or PDF is where it becomes a habit.

Profilbild von hamjji
hamjjivor 1 Monat

Naming Chinese and English as primary and the rest behind them is more honest than most releases. How wide is the gap between those tiers at 170M? Asking because the open source model I use leaked untranslated lines into dubs, and that was an alignment bug, not size.

Profilbild von Steven Casteel
Steven Casteelvor 1 Monat

The Japanese one sounded cute.

Profilbild von 安好
安好vor 23 Tagen

很棒,我要来试试

Profilbild von 푸른알약
푸른알약vor 1 Monat

한국어 쪽에서 이질감을 줄이고 싶다면 성조가 없는(또는 약한) 언어를 구분하고 학습시켜야 할 것 같네요. 한국어 표준말 측면에서는 상당히 이질적으로 들립니다. (거의 모든 원어민이 AI음성이라는 걸 바로 알아챌 정도랄까요?)

Profilbild von Harold Finch
Harold Finchvor 1 Monat

How does this fare against miso TTS?

Ähnliche Videos