Video yükleniyor...
Video Yüklenemedi
We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. 📚 Paper: 💻 Code: 📦 Model:
87,719 görüntüleme • 1 yıl önce •via X (Twitter)
11 Yorum

We also provide an online demo: Welcome to use and provide suggestions!

MMaDA is revolutionizing the multimodal game! I'm eager to see how its unified architecture and RL algorithms elevate text-to-image generation. This could be a real milestone in AI innovation. 🚀

That looks pretty sequential, isn’t the upside with diffusion to predict all the tokens every step?

We use unified diffusion modeling for training and provide flexible sampling strategies (Semi-AR or Non-AR) for better performance-speed trade-off

MMaDA is all you need.

Pls stop, we can’t keep up

CoT training data makes sense for autoregressive models, but what's the intuition for using it with a diffusion model?

Sounds like a game changer for bridging text and visuals 👏

not good prompt following.

That's a huge step towards truly unified multimodal AI

Curiosity piqued: does the model argue back? Asking for a friend. #AIorNot #DeepThoughtsNoDeepSleep #CoderInDisguise
