Загрузка видео...
Не удалось загрузить видео
We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. 📚 Paper: 💻 Code: 📦 Model:
87,719 просмотров • 1 год назад •via X (Twitter)
Комментарии: 11

We also provide an online demo: Welcome to use and provide suggestions!

MMaDA is revolutionizing the multimodal game! I'm eager to see how its unified architecture and RL algorithms elevate text-to-image generation. This could be a real milestone in AI innovation. 🚀

That looks pretty sequential, isn’t the upside with diffusion to predict all the tokens every step?

We use unified diffusion modeling for training and provide flexible sampling strategies (Semi-AR or Non-AR) for better performance-speed trade-off

MMaDA is all you need.

Pls stop, we can’t keep up

CoT training data makes sense for autoregressive models, but what's the intuition for using it with a diffusion model?

Sounds like a game changer for bridging text and visuals 👏

not good prompt following.

That's a huge step towards truly unified multimodal AI

Curiosity piqued: does the model argue back? Asking for a friend. #AIorNot #DeepThoughtsNoDeepSleep #CoderInDisguise
