Video wird geladen...
Video konnte nicht geladen werden
End to End Speech models are on fire - LLAMA-OMNI 8B - Apache licensed! 🔥 > Speech Encoder - Whisper Large v3 > LLM backbone - Llama 3.1 8B Instruct > Speech Decoder - HuBERT (UnitY) > Simultaneously generate Speech + Text > Less than 250 ms latency >... show more
47,921 Aufrufe • vor 1 Jahr •via X (Twitter)
10 Kommentare

Vaibhav (VB) Srivastavvor 1 Jahr
Model checkpoint:

Vaibhav (VB) Srivastavvor 1 Jahr
Github repo:

Qingkai Fangvor 1 Jahr
Thanks for sharing our work!

Vaibhav (VB) Srivastavvor 1 Jahr
🔥

Tommy D. Rossivor 1 Jahr
I wouldn't call this end to end, let's keep that term for single multi modal models that do everything by themselves

ThisAndThatvor 1 Jahr
less than 250ms latency on what?

Vaibhav (VB) Srivastavvor 1 Jahr
Time to first audio chunk according to their GH.

Waifuologyvor 1 Jahr
License looks good, but the voice quality isn't really there yet.

Hirovor 1 Jahr
Do you know what are supported languages?

Trying my best :-)vor 1 Jahr
Can it detect emotion?
