Video yükleniyor...
Video Yüklenemedi
Introducing Omni, one unified model can support any-to-any multimodal modeling, including multimodal understanding, image/video generation and editing, world modeling and 3D reconstruction. All in one that adopts standard mixture-of-experts arch with only 3B activations.
32,674 görüntüleme • 5 ay önce •via X (Twitter)
9 Yorum

Any-to-any training enables multimodal context unrolling where the model explicitly reasons across multiple modal representations before producing predictions. This reframes “unification” as a mechanism that scales context—not only in length, but in structure and utility for downstream decisions.

Started last summer, but shared much later and less detailed than we hoped. It was built by a small team under many constraints, with plenty of imperfections, but also with a lot of support from people who truely believed in it. It only made us believe more: the future is multimodal.

Homepage: ArXiv:

Congratulations Ceyuan! That’s great work!

Congrats Ceyuan!

Any GitHub repo link @CeyuanY for testing results

Congrats

interesting

Congrats🥳!

