Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Let's remove VAEs and ViTs from video models! 🚀 Introducing 𝗣𝗶𝘅𝗲𝗹𝗨𝗠𝗠: an encoder-free unified multimodal model for image and video understanding and generation, directly in pixel space. Code and model available today! 🌐 📄

71,377 görüntüleme • 1 gün önce •via X (Twitter)

19 Yorum

Cong Wei profil fotoğrafı
Cong Wei1 gün önce

We start from Qwen3-8B and turn it into a Mixture-of-Transformers. 🧩 Images → 16×16 patches, videos → 16x16x4 tubes, each via one linear layer. No VAE, no ViT. 🧠Separate understanding and generation experts share one self-attention over text, clean and noisy pixels. (2/N)

Cong Wei profil fotoğrafı
Cong Wei1 gün önce

PixelUMM achieves performance comparable to open-source baselines across image and video understanding and generation. (3/N)

Cong Wei profil fotoğrafı
Cong Wei1 gün önce

More demos available 🌐 (4/N)

Cong Wei profil fotoğrafı
Cong Wei1 gün önce

Joint work with an amazing team @nvidia : @xuanchi13 @HuanLing6 @huangjh_hjh @lealtaixe @FidlerSanja @zianwang97 @jayzhangjiewu @wmren993 @WenhuChen

Emily profil fotoğrafı
Emily1 gün önce

Hopefully, this approach will work out 🤞

Alex Ai | Tools & Updates profil fotoğrafı
Alex Ai | Tools & Updates1 gün önce

Encoder free pixel space approach is bold

Dongli Xu profil fotoğrafı
Dongli Xu1 gün önce

Wonderful work, next step is taking the diffusion head out. Full pixel model will win.

Arnas Uselis profil fotoğrafı
Arnas Uselis20 saat önce

Maybe let's not remove it? Isn't the context length already suffocating from all the frames?

Zara Tech & Tools profil fotoğrafı
Zara Tech & Tools1 gün önce

Exciting unified model approach

underscore advait patel profil fotoğrafı
underscore advait patel1 gün önce

I wonder if training a single MoE, similar to “Beyond LLMs” from Meta, would work better? Obviously you can’t start from Queen as easily, but they showed that the experts that are beneficial for vision understanding also typically get used for visual generation. It would be good interesting to see if this holds for video. Additionally, using MoE allows the model to learn how much capacity to allocate to each modality instead of having it be set ahead of time

Yueyan Li profil fotoğrafı
Yueyan Li1 gün önce

Great work!

gpu go brr... profil fotoğrafı
gpu go brr...1 gün önce

It seems like you should always use some kind of compressed encoding just to save on compute and memory. Even svd coeffients are ok.

Crio Songo profil fotoğrafı
Crio Songo1 gün önce

No VAE/ViT, direct pixel space—curious how this scales to long videos.

Vera Ai | Tools & Updates profil fotoğrafı
Vera Ai | Tools & Updates1 gün önce

Direct pixel-space multimodal generation is a fascinating approach to video models

Solomon profil fotoğrafı
Solomon1 gün önce

The F8 ablation is a fun surprise: 4-FPS tubelets and 1-FPS frames land very close on all four video benchmarks. Have you looked at fast-motion or temporal-order subsets to see where the denser interface helps?

ONE ☗ profil fotoğrafı
ONE ☗1 gün önce

Excited about it!

Let’s goo profil fotoğrafı
Let’s goo1 gün önce

Encoder-free pixel space unification bold move impressive simplification concept indeed

PookieNumnums profil fotoğrafı
PookieNumnums1 gün önce

Interesting but what's the point or benefit. The motion and consistency seems a bit lacking. Any plans to do this with larger models?

Veyra | AI & Tech profil fotoğrafı
Veyra | AI & Tech1 gün önce

Encoder-free unified model operating directly in pixel space wild.

Benzer Videolar