正在加载视频...
视频加载失败
🚀 Introducing Emu3.5 — a large-scale multimodal world model that natively predicts the next vision-language state. 🔥 Trained on over 10T interleaved vision-language tokens and enhanced with reinforcement learning, Emu3.5 achieves powerful multimodal reasoning and generation. ⚡ Powered by our new Discrete Diffusion Adaptation (DiDA) for 20× faster inference.... show more
51,880 次观看 • 9 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
