Video yükleniyor...
Video Yüklenemedi
Let's remove VAEs and ViTs from video models! 🚀 Introducing 𝗣𝗶𝘅𝗲𝗹𝗨𝗠𝗠: an encoder-free unified multimodal model for image and video understanding and generation, directly in pixel space. Code and model available today! 🌐 📄
71,377 görüntüleme • 1 gün önce •via X (Twitter)
19 Yorum

We start from Qwen3-8B and turn it into a Mixture-of-Transformers. 🧩 Images → 16×16 patches, videos → 16x16x4 tubes, each via one linear layer. No VAE, no ViT. 🧠Separate understanding and generation experts share one self-attention over text, clean and noisy pixels. (2/N)

PixelUMM achieves performance comparable to open-source baselines across image and video understanding and generation. (3/N)

More demos available 🌐 (4/N)

Joint work with an amazing team @nvidia : @xuanchi13 @HuanLing6 @huangjh_hjh @lealtaixe @FidlerSanja @zianwang97 @jayzhangjiewu @wmren993 @WenhuChen

Hopefully, this approach will work out 🤞

Encoder free pixel space approach is bold

Wonderful work, next step is taking the diffusion head out. Full pixel model will win.

Maybe let's not remove it? Isn't the context length already suffocating from all the frames?

Exciting unified model approach

I wonder if training a single MoE, similar to “Beyond LLMs” from Meta, would work better? Obviously you can’t start from Queen as easily, but they showed that the experts that are beneficial for vision understanding also typically get used for visual generation. It would be good interesting to see if this holds for video. Additionally, using MoE allows the model to learn how much capacity to allocate to each modality instead of having it be set ahead of time

Great work!

It seems like you should always use some kind of compressed encoding just to save on compute and memory. Even svd coeffients are ok.

No VAE/ViT, direct pixel space—curious how this scales to long videos.

Direct pixel-space multimodal generation is a fascinating approach to video models

The F8 ablation is a fun surprise: 4-FPS tubelets and 1-FPS frames land very close on all four video benchmarks. Have you looked at fast-motion or temporal-order subsets to see where the denser interface helps?

Excited about it!

Encoder-free pixel space unification bold move impressive simplification concept indeed

Interesting but what's the point or benefit. The motion and consistency seems a bit lacking. Any plans to do this with larger models?

Encoder-free unified model operating directly in pixel space wild.
