Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

The world’s smallest Transformer-based TTS model? We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.

74,884 görüntüleme • 1 ay önce •via X (Twitter)

39 Yorum

Samuel Zeng profil fotoğrafı
Samuel Zeng1 ay önce

What can a 0.1B-class TTS model actually sound like? Listen to the voiceover in this demo video. Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian, Japanese, and Korean Small in scale. Surprisingly capable in speech.

Samuel Zeng profil fotoğrafı
Samuel Zeng1 ay önce

On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves: • English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2 These results are achieved with an approximately 170M-parameter main model.

Samuel Zeng profil fotoğrafı
Samuel Zeng1 ay önce

The model, codec, tokenizer, processor, and inference code are now available: Model: Try it, test the voice cloning capability, and share your feedback.

Samuel Zeng profil fotoğrafı
Samuel Zeng26 gün önce

update: INT8 version only 170MB

Desh Raj profil fotoğrafı
Desh Raj1 ay önce

Nice work, but I think Pocket-TTS ( is ~100M params.

Samuel Zeng profil fotoğrafı
Samuel Zeng1 ay önce

Thanks. Good catch — Pocket-TTS is indeed smaller at ~100M parameters. Our focus is compact multilingual TTS, zero-shot voice cloning, and efficient end-to-end speech generation.

RAZA | AI EXPLORER profil fotoğrafı
RAZA | AI EXPLORER1 ay önce

Tiny footprint, impressive voice quality. Open-source TTS keeps getting more exciting.

Moez Zhioua profil fotoğrafı
Moez Zhioua1 ay önce

170M parameters and sub 2% English WER is wild. Shows multilingual voice cloning can be high quality without a massive model. Size isn’t the only factor for TTS performance

Romain profil fotoğrafı
Romain1 ay önce

It is possible to run it on mobile like Supertonic ?

Samuel Zeng profil fotoğrafı
Samuel Zeng1 ay önce

Sure, It works.

Habtamu Asefa profil fotoğrafı
Habtamu Asefa1 ay önce

This is insane 😲

Ha Hoang profil fotoğrafı
Ha Hoang1 ay önce

huh wow ~0.2b for this quality?

Math3matica profil fotoğrafı
Math3matica1 ay önce

Kokoro is 82 million. Still excited to check this out!

Internet Labs profil fotoğrafı
Internet Labs1 ay önce

ElevenLabs executives watching a 170M parameter open-source model do zero-shot voice cloning locally for free: 👁️👄👁️

luis profil fotoğrafı
luis1 ay önce

@Presidentlin

David Orban profil fotoğrafı
David Orban1 ay önce

This is the useful shape of an exponential: not a bigger model, a closed loop! Speech that used to need a cluster now fits in 170 million parameters. The compounding is local. The next question is, which other loops have already closed this way and still get talked about as if they need the frontier?

Forlais profil fotoğrafı
Forlais1 ay önce

Fair, 170M holding cloud-level voice quality says the ceiling was never really about size, a point we came at from the other side with a floatless architecture rather than a traditional transformer.

Piyush profil fotoğrafı
Piyush1 ay önce

@Petucato

🦀ꃳꁲꌅꋖ🦀 profil fotoğrafı
🦀ꃳꁲꌅꋖ🦀1 ay önce

Będę pod wrażeniem jak dodacie Polski. Kyutai ma najmniejsze modele. Nikt nie pamięta o języku polskim.

Mahimai Raja J ‎ profil fotoğrafı
Mahimai Raja J ‎1 ay önce

love it!

kfant profil fotoğrafı
kfant1 ay önce

impressive lets see if my hermes agent can talk to me in squeeky anime voice with Audio8

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

170M params, cloud-level quality, impressive.

Human Being profil fotoğrafı
Human Being1 ay önce

This is incredible. I was just searching for smaller cloning TTS models today and coincidentally found this now.

Eva Morgan profil fotoğrafı
Eva Morgan1 ay önce

170M parameters with multilingual TTS and zero shot voice cloning is impressive. The smaller footprint could make local deployment much more practical.

Johnny Depp profil fotoğrafı
Johnny Depp1 ay önce

Wow how about phone version?

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan1 ay önce

170m is compelling, but multilingual pronunciation drift matters more than demo voice quality

cordivai | Machine Learning & AI profil fotoğrafı
cordivai | Machine Learning & AI1 ay önce

Useful for students implementing papers too. Before chasing the newest method, build a simple baseline, reproduce the core setting, then add ablations so the improvement is actually explainable.

Temel Günaydın profil fotoğrafı
Temel Günaydın1 ay önce

Awesome 👏

Rithul profil fotoğrafı
Rithul1 ay önce

Fi fa fi fi faaa

Mattias Johansson profil fotoğrafı
Mattias Johansson1 ay önce

Did you change the license to non-commercial?

rusa profil fotoğrafı
rusa1 ay önce

at 0.1B the interesting constraint is prosody. tiny tts models nail phonemes but struggle with long-range intonation, the sentence-level melody. if you kept it natural at that size that's the real achievement, since the payoff is running fully on-device with no latency.

Beeno Tung profil fotoğrafı
Beeno Tung1 ay önce

I did a small scale testing, so far the best accuracy for cantonuese is sensevoice and whisper cantonese fine tune, audio8 is close, better than whisper base model

Michael profil fotoğrafı
Michael28 gün önce

looks and sounds great, isn't kokorro tts 88million parameters though?

ThuNoJoy | SayTXT for ADHD Reading profil fotoğrafı
ThuNoJoy | SayTXT for ADHD Reading1 ay önce

The small-footprint part matters a lot for reading tools. Natural TTS is great, but fast enough to use on every article or PDF is where it becomes a habit.

hamjji profil fotoğrafı
hamjji1 ay önce

Naming Chinese and English as primary and the rest behind them is more honest than most releases. How wide is the gap between those tiers at 170M? Asking because the open source model I use leaked untranslated lines into dubs, and that was an alignment bug, not size.

Steven Casteel profil fotoğrafı
Steven Casteel1 ay önce

The Japanese one sounded cute.

安好 profil fotoğrafı
安好22 gün önce

很棒,我要来试试

푸른알약 profil fotoğrafı
푸른알약1 ay önce

한국어 쪽에서 이질감을 줄이고 싶다면 성조가 없는(또는 약한) 언어를 구분하고 학습시켜야 할 것 같네요. 한국어 표준말 측면에서는 상당히 이질적으로 들립니다. (거의 모든 원어민이 AI음성이라는 걸 바로 알아챌 정도랄까요?)

Harold Finch profil fotoğrafı
Harold Finch1 ay önce

How does this fare against miso TTS?

Benzer Videolar