Загрузка видео...

Не удалось загрузить видео

На главную

We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. 📚 Paper: 💻 Code: 📦 Model:

87,719 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля Ling Yang
Ling Yang1 год назад

We also provide an online demo: Welcome to use and provide suggestions!

Фото профиля FemMiraMancer
FemMiraMancer1 год назад

MMaDA is revolutionizing the multimodal game! I'm eager to see how its unified architecture and RL algorithms elevate text-to-image generation. This could be a real milestone in AI innovation. 🚀

Фото профиля Stellan Haglund
Stellan Haglund1 год назад

That looks pretty sequential, isn’t the upside with diffusion to predict all the tokens every step?

Фото профиля Ling Yang
Ling Yang1 год назад

We use unified diffusion modeling for training and provide flexible sampling strategies (Semi-AR or Non-AR) for better performance-speed trade-off

Фото профиля Ruben Tous
Ruben Tous1 год назад

MMaDA is all you need.

Фото профиля Mr. John Smith
Mr. John Smith1 год назад

Pls stop, we can’t keep up

Фото профиля FluffyCat
FluffyCat1 год назад

CoT training data makes sense for autoregressive models, but what's the intuition for using it with a diffusion model?

Фото профиля Hyperstack
Hyperstack1 год назад

Sounds like a game changer for bridging text and visuals 👏

Фото профиля Bali as a Colony of Jakarta
Bali as a Colony of Jakarta1 год назад

not good prompt following.

Фото профиля Sam Woods
Sam Woods1 год назад

That's a huge step towards truly unified multimodal AI

Фото профиля MangoMagic™️
MangoMagic™️1 год назад

Curiosity piqued: does the model argue back? Asking for a friend. #AIorNot #DeepThoughtsNoDeepSleep #CoderInDisguise

Похожие видео

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 просмотров • 1 год назад