Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. 📚 Paper: 💻 Code: 📦 Model:

87,719 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von Ling Yang
Ling Yangvor 1 Jahr

We also provide an online demo: Welcome to use and provide suggestions!

Profilbild von FemMiraMancer
FemMiraMancervor 1 Jahr

MMaDA is revolutionizing the multimodal game! I'm eager to see how its unified architecture and RL algorithms elevate text-to-image generation. This could be a real milestone in AI innovation. 🚀

Profilbild von Stellan Haglund
Stellan Haglundvor 1 Jahr

That looks pretty sequential, isn’t the upside with diffusion to predict all the tokens every step?

Profilbild von Ling Yang
Ling Yangvor 1 Jahr

We use unified diffusion modeling for training and provide flexible sampling strategies (Semi-AR or Non-AR) for better performance-speed trade-off

Profilbild von Ruben Tous
Ruben Tousvor 1 Jahr

MMaDA is all you need.

Profilbild von Mr. John Smith
Mr. John Smithvor 1 Jahr

Pls stop, we can’t keep up

Profilbild von FluffyCat
FluffyCatvor 1 Jahr

CoT training data makes sense for autoregressive models, but what's the intuition for using it with a diffusion model?

Profilbild von Hyperstack
Hyperstackvor 1 Jahr

Sounds like a game changer for bridging text and visuals 👏

Profilbild von Bali as a Colony of Jakarta
Bali as a Colony of Jakartavor 1 Jahr

not good prompt following.

Profilbild von Sam Woods
Sam Woodsvor 1 Jahr

That's a huge step towards truly unified multimodal AI

Profilbild von MangoMagic™️
MangoMagic™️vor 1 Jahr

Curiosity piqued: does the model argue back? Asking for a friend. #AIorNot #DeepThoughtsNoDeepSleep #CoderInDisguise

Ähnliche Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 Aufrufe • vor 1 Jahr