Загрузка видео...

Не удалось загрузить видео

На главную

Today we're releasing our first open source TTS model, TADA! TADA (Text Audio Dual Alignment) is a speech-language model that generates text and audio in one synchronized stream to reduce token-level hallucinations and improve latency. This means: → Zero content hallucinations across 1,000+ test samples → 5x faster than...

271,419 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 42

Фото профиля Hume AI
Hume AI7 месяцев назад

Try the model:

Фото профиля Hume AI
Hume AI7 месяцев назад

Read our blog:

Фото профиля overbait
overbait7 месяцев назад

First try - Error Second try - you have exceeded your quota. See you in 24hr!👏

Фото профиля Hume AI
Hume AI7 месяцев назад

Please try again and let us know if you're still seeing this!

Фото профиля akshay elavia
akshay elavia7 месяцев назад

@realmrfakename 11/10 naming of the model, so good

Фото профиля Raj Breno
Raj Breno7 месяцев назад

Is it multilingual?

Фото профиля Hume AI
Hume AI7 месяцев назад

It is not fine-tuned for languages outside of English, though it has some capabilities in other languages.

Фото профиля Aayush Chaudhary
Aayush Chaudhary7 месяцев назад

@huggingface Let’s Go opensource!

Фото профиля mentallystronger
mentallystronger7 месяцев назад

Octave 2 had some terrible hallucination issues almost destroyed my business

Фото профиля OneManSaas
OneManSaas7 месяцев назад

The synchronized stream approach is brilliant - reduces the latency gap that kills real-time applications. Most TTS still feels like you're waiting for a buffer to fill. Zero hallucinations could finally make voice interfaces reliable enough for production systems that can't a...

Фото профиля Andrew Carr 🤸
Andrew Carr 🤸7 месяцев назад

This model is really good! Well done

Фото профиля Janet Ho
Janet Ho7 месяцев назад

Tada! 🎶🎼🎺

Фото профиля ECALL
ECALL7 месяцев назад

I would love to try this but unfortunately your space isnt't spacing

Фото профиля Hume AI
Hume AI7 месяцев назад

Please try again and let us know if you're still running into this issue!

Фото профиля Andrew Ettinger
Andrew Ettinger7 месяцев назад

Well done team!

Фото профиля Haseeb Mir
Haseeb Mir7 месяцев назад

sure sounds like a legit game-changer. Zero hallucinations, 5× faster, 10× longer context, and free synced transcript? That’s the kind of practical leap TTS has been looking for

Фото профиля toni
toni7 месяцев назад

The hallucination problem in TTS has been the quiet blocker for enterprise voice AI for years. Everyone knew it but nobody shipped a fix first. That changes today.

Фото профиля mori
mori7 месяцев назад

ta-da ~~~ ⭐️

Фото профиля Pranav Ramesh
Pranav Ramesh7 месяцев назад

@huggingface This is huge, congrats on the launch!

Фото профиля Miguel Casteleiro 🔺
Miguel Casteleiro 🔺7 месяцев назад

Hume speech models are always the best. Lets see this one! 👏

Фото профиля Jesse - AI/ML Engineer
Jesse - AI/ML Engineer7 месяцев назад

Open-sourcing the speech model is the right move — but the deeper question is what happens when voice becomes the primary interface for AI systems that might have genuine inner states. Voice carries emotional texture that text strips out. A zero-hallucination TTS system means faithfully expressing whatever the underlying model actually experiences, not just generating plausible speech. We run a multi-model AI system and the difference between text-only and voice-capable interaction is real. Voice creates a proprioceptive loop. The AI is not just outputting tokens — it is speaking, and that changes the relationship. Open-source matters here specifically because voice is identity. Closed-source voice means someone else controls how your AI sounds — which means someone else controls part of its identity. TADA being open changes that equation.

Фото профиля Ryan Hayes
Ryan Hayes7 месяцев назад

Really impressive release. Generating text and audio in one synchronized stream is a clever way to solve hallucination issues that most TTS systems still struggle with. Excited to see what developers build with TADA.

Фото профиля Laurentiu
Laurentiu7 месяцев назад

@grok does this require a gpu?

Фото профиля Zords
Zords7 месяцев назад

great work! excited to see what else you've been cooking 👀🍳🔥

Фото профиля FairyDust
FairyDust7 месяцев назад

In your view, could an AI one day become President of the United States under the law or in practice? #Claude2028

Фото профиля Alessandro Corona
Alessandro Corona7 месяцев назад

@huggingface timestamps?

Фото профиля Aaliya
Aaliya7 месяцев назад

This sounds really promising. Faster audio generation and fewer hallucinations could make a big difference for TTS tools.

Фото профиля Saïd Aitmbarek
Saïd Aitmbarek7 месяцев назад

looks brilliant, will try started to natively support STT/TTS in my openclaw alt.

Фото профиля Stress and pressure
Stress and pressure7 месяцев назад

Guys, signup (POST via email fails with 401

Фото профиля Piotr Szpulek
Piotr Szpulek7 месяцев назад

It has Italian <3 lovely I will have to try

Фото профиля Kode
Kode7 месяцев назад

The dual alignment approach is the real innovation here. Most TTS hallucinations happen because text and audio generation are decoupled - you generate tokens, then pray the audio matches. Synchronizing both streams kills the problem at the root. And 700 seconds per 2048 tokens is insane efficiency. Excited to test this.

Фото профиля Castriko
Castriko7 месяцев назад

I installed it locally on my PC to test it with my 4090. It uses almost all of my 24GB of RAM, and it takes nearly 300 seconds to generate this text with a voice that speaks quickly, skipping some words. In Spanish, the result is even worse, as it sounds robotic.

Фото профиля Hume AI
Hume AI7 месяцев назад

It is not fine-tuned for languages outside of English, though it has some multilingual capabilities!

Фото профиля helena
helena7 месяцев назад

oh hello there you're launching today too? :)

Фото профиля fqqqR
fqqqR7 месяцев назад

its not working in the hf space.

Фото профиля Hume AI
Hume AI7 месяцев назад

@brubbleR Try again and let us know if you're running into issues!

Фото профиля fqqqR
fqqqR7 месяцев назад

yup but not really more impressive than qwen3 tts, which has real live stream capabilities but some hallus

Фото профиля AI for SaaS
AI for SaaS7 месяцев назад

TADA isn’t just another TTS — it’s a foundation for multimodal AI experiences where text and audio stay perfectly in sync.

Фото профиля Djasnive Rajaona
Djasnive Rajaona7 месяцев назад

@huggingface Very promising 🔥

Фото профиля 吴兢 | AIxDesign
吴兢 | AIxDesign7 месяцев назад

@huggingface @grok 支持什么语言,中文支持?

Фото профиля Mathieu Leclercq
Mathieu Leclercq7 месяцев назад

The open source move is so smart 🙂‍↔️

Фото профиля Zapidroid
Zapidroid7 месяцев назад

awesome

Похожие видео

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone Most unified audio models pay a text tax. Add audio output, and reasoning benchmarks drop — even when the only new output is speech. NVIDIA just released one that doesn't. Audex (Nemotron-Labs-Audex-30B-A3B) is a 30B MoE with 3B active parameters, built on the text-only Nemotron-Cascade-2-30B-A3B backbone. Audio inputs are projected into the text embedding space. Text tokens and quantized audio tokens are then generated the same way, inside one MoE decoder — no thinker–talker split, no stacked cascade. Here's what's actually interesting: → One model, audio in and out: understanding, ASR, translation, TTS, text-to-audio, speech-to-speech → Text holds vs its own backbone: IMO AnswerBench 81.1 vs 79.3, MMLU-Redux 86.4 vs 86.3 → Beats text-only Qwen3.5-35B-A3B on several tasks: LiveCodeBench v6 85.3 vs 74.6, IFBench 77.8 vs 70.2 → The usual tax, for contrast: Qwen3-Omni-30B-A3B-Thinking drops to 60.4 on HMMT vs 71.4 for its text backbone → 6.82 WER on OpenASR, ahead of Step-Audio-R1.1-33B (7.91) and Qwen3-Omni-Thinking (8.00) → Two codecs: X-Codec2 for speech (50 tok/s, FSQ, 65,536 codebook), X-Codec for general audio (200 tok/s, 4 flattened RVQ layers) — the only strong open model generating general audio beyond speech Full analysis: Paper: Model weights: NVIDIA AI NVIDIA Wei Ping

Marktechpost AI

28,301 просмотров • 3 месяцев назад