ๆญฃๅจๅ ่ฝฝ่ง้ข...
่ง้ขๅ ่ฝฝๅคฑ่ดฅ
End to End Speech models are on fire - LLAMA-OMNI 8B - Apache licensed! ๐ฅ > Speech Encoder - Whisper Large v3 > LLM backbone - Llama 3.1 8B Instruct > Speech Decoder - HuBERT (UnitY) > Simultaneously generate Speech + Text > Less than 250 ms latency >... show more
47,921 ๆฌก่ง็ โข 1 ๅนดๅ โขvia X (Twitter)
10 ๆก่ฏ่ฎบ

Vaibhav (VB) Srivastav1 ๅนดๅ
Model checkpoint:

Vaibhav (VB) Srivastav1 ๅนดๅ
Github repo:

Qingkai Fang1 ๅนดๅ
Thanks for sharing our work!

Vaibhav (VB) Srivastav1 ๅนดๅ
๐ฅ

Tommy D. Rossi1 ๅนดๅ
I wouldn't call this end to end, let's keep that term for single multi modal models that do everything by themselves

ThisAndThat1 ๅนดๅ
less than 250ms latency on what?

Vaibhav (VB) Srivastav1 ๅนดๅ
Time to first audio chunk according to their GH.

Waifuology1 ๅนดๅ
License looks good, but the voice quality isn't really there yet.

Hiro1 ๅนดๅ
Do you know what are supported languages?

Trying my best :-)1 ๅนดๅ
Can it detect emotion?
