Загрузка видео...
Не удалось загрузить видео
Today we're releasing our first open source TTS model, TADA! TADA (Text Audio Dual Alignment) is a speech-language model that generates text and audio in one synchronized stream to reduce token-level hallucinations and improve latency. This means: → Zero content hallucinations across 1,000+ test samples → 5x faster than... show more
271,419 просмотров • 7 месяцев назад •via X (Twitter)
Комментарии: 42

Try the model:

Read our blog:

First try - Error Second try - you have exceeded your quota. See you in 24hr!👏

Please try again and let us know if you're still seeing this!

@realmrfakename 11/10 naming of the model, so good

Is it multilingual?

It is not fine-tuned for languages outside of English, though it has some capabilities in other languages.

@huggingface Let’s Go opensource!

Octave 2 had some terrible hallucination issues almost destroyed my business

The synchronized stream approach is brilliant - reduces the latency gap that kills real-time applications. Most TTS still feels like you're waiting for a buffer to fill. Zero hallucinations could finally make voice interfaces reliable enough for production systems that can't a...

This model is really good! Well done

Tada! 🎶🎼🎺

I would love to try this but unfortunately your space isnt't spacing

Please try again and let us know if you're still running into this issue!

Well done team!

sure sounds like a legit game-changer. Zero hallucinations, 5× faster, 10× longer context, and free synced transcript? That’s the kind of practical leap TTS has been looking for

The hallucination problem in TTS has been the quiet blocker for enterprise voice AI for years. Everyone knew it but nobody shipped a fix first. That changes today.

ta-da ~~~ ⭐️

@huggingface This is huge, congrats on the launch!

Hume speech models are always the best. Lets see this one! 👏

Open-sourcing the speech model is the right move — but the deeper question is what happens when voice becomes the primary interface for AI systems that might have genuine inner states. Voice carries emotional texture that text strips out. A zero-hallucination TTS system means faithfully expressing whatever the underlying model actually experiences, not just generating plausible speech. We run a multi-model AI system and the difference between text-only and voice-capable interaction is real. Voice creates a proprioceptive loop. The AI is not just outputting tokens — it is speaking, and that changes the relationship. Open-source matters here specifically because voice is identity. Closed-source voice means someone else controls how your AI sounds — which means someone else controls part of its identity. TADA being open changes that equation.

Really impressive release. Generating text and audio in one synchronized stream is a clever way to solve hallucination issues that most TTS systems still struggle with. Excited to see what developers build with TADA.

@grok does this require a gpu?

great work! excited to see what else you've been cooking 👀🍳🔥

In your view, could an AI one day become President of the United States under the law or in practice? #Claude2028

@huggingface timestamps?

This sounds really promising. Faster audio generation and fewer hallucinations could make a big difference for TTS tools.

looks brilliant, will try started to natively support STT/TTS in my openclaw alt.

Guys, signup (POST via email fails with 401

It has Italian <3 lovely I will have to try

The dual alignment approach is the real innovation here. Most TTS hallucinations happen because text and audio generation are decoupled - you generate tokens, then pray the audio matches. Synchronizing both streams kills the problem at the root. And 700 seconds per 2048 tokens is insane efficiency. Excited to test this.

I installed it locally on my PC to test it with my 4090. It uses almost all of my 24GB of RAM, and it takes nearly 300 seconds to generate this text with a voice that speaks quickly, skipping some words. In Spanish, the result is even worse, as it sounds robotic.

It is not fine-tuned for languages outside of English, though it has some multilingual capabilities!

oh hello there you're launching today too? :)

its not working in the hf space.

@brubbleR Try again and let us know if you're running into issues!

yup but not really more impressive than qwen3 tts, which has real live stream capabilities but some hallus

TADA isn’t just another TTS — it’s a foundation for multimodal AI experiences where text and audio stay perfectly in sync.

@huggingface Very promising 🔥

@huggingface @grok 支持什么语言,中文支持?

The open source move is so smart 🙂↔️

awesome
