Загрузка видео...

Не удалось загрузить видео

На главную

We just solved text-to-speech AI. This model can simulate perfect emotion, screaming and show genuine alarm. — clearly beats 11 labs and Sesame — it’s only 1.6B params — streams realtime on 1 GPU — made by a 1.5 person team in Korea!! It's called Dia by Nari Labs.

711,414 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля Deedy
Deedy1 год назад

Source:

Фото профиля Deedy
Deedy1 год назад

The future is about to look really weird. Audio may have just crossed the uncanny valley (like parts of text and Ike have) into most-humans-wont-know-this-is-AI territory

Фото профиля MightyBot
MightyBot1 год назад

🧠 Unified Search. Smarter Meetings. Effortless CRM. MightyBot is your AI agent platform for seamless workflows—record meetings, automate CRM updates, and find answers across apps in seconds. 🌟 Focus on what matters. We'll handle the grind.

Фото профиля Yuchen Jin
Yuchen Jin1 год назад

what is the 0.5 person in the 1.5 person team? 😂

Фото профиля Deedy
Deedy1 год назад

Part time research engineer!

Фото профиля Mudit Juneja
Mudit Juneja1 год назад

Who are we here? Are you tied to this project?

Фото профиля Deedy
Deedy1 год назад

We = humanity

Фото профиля Cr33d
Cr33d1 год назад

1.5 people?! Did the 0.5 person just handle the screaming?

Фото профиля Rithik Chopra
Rithik Chopra1 год назад

Damn that’s crazy!!!

Фото профиля Albert Sebastian
Albert Sebastian1 год назад

whats your take on hume ai?

Фото профиля Cr33d
Cr33d1 год назад

Perfect emotion? Finally, my toaster can apologize for burning my toast! 😂

Похожие видео

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone Most unified audio models pay a text tax. Add audio output, and reasoning benchmarks drop — even when the only new output is speech. NVIDIA just released one that doesn't. Audex (Nemotron-Labs-Audex-30B-A3B) is a 30B MoE with 3B active parameters, built on the text-only Nemotron-Cascade-2-30B-A3B backbone. Audio inputs are projected into the text embedding space. Text tokens and quantized audio tokens are then generated the same way, inside one MoE decoder — no thinker–talker split, no stacked cascade. Here's what's actually interesting: → One model, audio in and out: understanding, ASR, translation, TTS, text-to-audio, speech-to-speech → Text holds vs its own backbone: IMO AnswerBench 81.1 vs 79.3, MMLU-Redux 86.4 vs 86.3 → Beats text-only Qwen3.5-35B-A3B on several tasks: LiveCodeBench v6 85.3 vs 74.6, IFBench 77.8 vs 70.2 → The usual tax, for contrast: Qwen3-Omni-30B-A3B-Thinking drops to 60.4 on HMMT vs 71.4 for its text backbone → 6.82 WER on OpenASR, ahead of Step-Audio-R1.1-33B (7.91) and Qwen3-Omni-Thinking (8.00) → Two codecs: X-Codec2 for speech (50 tok/s, FSQ, 65,536 codebook), X-Codec for general audio (200 tok/s, 4 flattened RVQ layers) — the only strong open model generating general audio beyond speech Full analysis: Paper: Model weights: NVIDIA AI NVIDIA Wei Ping

Marktechpost AI

28,301 просмотров • 2 месяцев назад