Загрузка видео...
Не удалось загрузить видео
Mini-Omni 2 understands image, audio and text inputs all via end-to-end voice conversations with users 🔥 > Understands and processes images, speech, and text > Generates real-time speech responses > Supports interruptions during speech Technical Overview: > Concatenates image, audio, and text features for input. > Uses text-guided delayed... show more
45,453 просмотров • 1 год назад •via X (Twitter)
Комментарии: 8

Vaibhav (VB) Srivastav1 год назад
Check out the model here:

Julien Blanchon 🇺🇦1 год назад
SNAC 🙏

Vaibhav (VB) Srivastav1 год назад
Hubert is the real 🐐

0xCrashX1 год назад
For some reason the " file is flagged as suspicious 🫤

Species ⚡️ 🌘 🌍1 год назад
Can it be used to train / fine turn for other languages than just English / Chinese?

Jay Sensei👾1 год назад
It's using whisper n Qwen2, any space available to try it? How much system gig needed to be run locally?

Godcreated1 год назад
Rise of the conversational cyborgs - MIT's gift to humanity.

ToneDice1 год назад
Wow could be great for blind people
