Загрузка видео...

Не удалось загрузить видео

На главную

Mini-Omni 2 understands image, audio and text inputs all via end-to-end voice conversations with users 🔥 > Understands and processes images, speech, and text > Generates real-time speech responses > Supports interruptions during speech Technical Overview: > Concatenates image, audio, and text features for input. > Uses text-guided delayed...

45,453 просмотров • 1 год назад •via X (Twitter)

Комментарии: 8

Фото профиля Vaibhav (VB) Srivastav
Vaibhav (VB) Srivastav1 год назад

Check out the model here:

Фото профиля Julien Blanchon 🇺🇦
Julien Blanchon 🇺🇦1 год назад

SNAC 🙏

Фото профиля Vaibhav (VB) Srivastav
Vaibhav (VB) Srivastav1 год назад

Hubert is the real 🐐

Фото профиля 0xCrashX
0xCrashX1 год назад

For some reason the " file is flagged as suspicious 🫤

Фото профиля Species ⚡️ 🌘 🌍
Species ⚡️ 🌘 🌍1 год назад

Can it be used to train / fine turn for other languages than just English / Chinese?

Фото профиля Jay Sensei👾
Jay Sensei👾1 год назад

It's using whisper n Qwen2, any space available to try it? How much system gig needed to be run locally?

Фото профиля Godcreated
Godcreated1 год назад

Rise of the conversational cyborgs - MIT's gift to humanity.

Фото профиля ToneDice
ToneDice1 год назад

Wow could be great for blind people

Похожие видео

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone Most unified audio models pay a text tax. Add audio output, and reasoning benchmarks drop — even when the only new output is speech. NVIDIA just released one that doesn't. Audex (Nemotron-Labs-Audex-30B-A3B) is a 30B MoE with 3B active parameters, built on the text-only Nemotron-Cascade-2-30B-A3B backbone. Audio inputs are projected into the text embedding space. Text tokens and quantized audio tokens are then generated the same way, inside one MoE decoder — no thinker–talker split, no stacked cascade. Here's what's actually interesting: → One model, audio in and out: understanding, ASR, translation, TTS, text-to-audio, speech-to-speech → Text holds vs its own backbone: IMO AnswerBench 81.1 vs 79.3, MMLU-Redux 86.4 vs 86.3 → Beats text-only Qwen3.5-35B-A3B on several tasks: LiveCodeBench v6 85.3 vs 74.6, IFBench 77.8 vs 70.2 → The usual tax, for contrast: Qwen3-Omni-30B-A3B-Thinking drops to 60.4 on HMMT vs 71.4 for its text backbone → 6.82 WER on OpenASR, ahead of Step-Audio-R1.1-33B (7.91) and Qwen3-Omni-Thinking (8.00) → Two codecs: X-Codec2 for speech (50 tok/s, FSQ, 65,536 codebook), X-Codec for general audio (200 tok/s, 4 flattened RVQ layers) — the only strong open model generating general audio beyond speech Full analysis: Paper: Model weights: NVIDIA AI NVIDIA Wei Ping

Marktechpost AI

28,301 просмотров • 2 месяцев назад