Loading video...

Video Failed to Load

Go Home

We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. 📚 Paper: 💻 Code: 📦 Model:

87,719 views • 1 year ago •via X (Twitter)

11 Comments

Ling Yang's profile picture
Ling Yang1 year ago

We also provide an online demo: Welcome to use and provide suggestions!

FemMiraMancer's profile picture
FemMiraMancer1 year ago

MMaDA is revolutionizing the multimodal game! I'm eager to see how its unified architecture and RL algorithms elevate text-to-image generation. This could be a real milestone in AI innovation. 🚀

Stellan Haglund's profile picture
Stellan Haglund1 year ago

That looks pretty sequential, isn’t the upside with diffusion to predict all the tokens every step?

Ling Yang's profile picture
Ling Yang1 year ago

We use unified diffusion modeling for training and provide flexible sampling strategies (Semi-AR or Non-AR) for better performance-speed trade-off

Ruben Tous's profile picture
Ruben Tous1 year ago

MMaDA is all you need.

Mr. John Smith's profile picture
Mr. John Smith1 year ago

Pls stop, we can’t keep up

FluffyCat's profile picture
FluffyCat1 year ago

CoT training data makes sense for autoregressive models, but what's the intuition for using it with a diffusion model?

Hyperstack's profile picture
Hyperstack1 year ago

Sounds like a game changer for bridging text and visuals 👏

Bali as a Colony of Jakarta's profile picture
Bali as a Colony of Jakarta1 year ago

not good prompt following.

Sam Woods's profile picture
Sam Woods1 year ago

That's a huge step towards truly unified multimodal AI

MangoMagic™️'s profile picture
MangoMagic™️1 year ago

Curiosity piqued: does the model argue back? Asking for a friend. #AIorNot #DeepThoughtsNoDeepSleep #CoderInDisguise

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 views • 1 year ago