Loading video...
Video Failed to Load
🚀 Introducing Emu3.5 — a large-scale multimodal world model that natively predicts the next vision-language state. 🔥 Trained on over 10T interleaved vision-language tokens and enhanced with reinforcement learning, Emu3.5 achieves powerful multimodal reasoning and generation. ⚡ Powered by our new Discrete Diffusion Adaptation (DiDA) for 20× faster inference.... show more
51,880 views • 9 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
