Video yükleniyor...
Video Yüklenemedi
Cosmos Policy Fine-Tuning Video Models for Visuomotor Control and Planning
10,355 görüntüleme • 8 ay önce •via X (Twitter)
3 Yorum

paper:

Thank you for sharing our work! 🙏

Video generation models aren’t just for making clips anymore—they’re becoming the robot brain itself. The part that really clicked for me in this NVIDIA/Stanford paper is the simplicity. Instead of bolting on complex control modules, they just treat robot actions as "latent frames" within the video sequence. Because of this, a single fine-tuned video model (Cosmos-Predict2) acts as the policy, the world simulator, and the value function all at once. It’s hitting 98.5% success on the LIBERO benchmark just by "dreaming" the right future frames and actions together. Makes you wonder—is the path to AGI just a really, really good video predictor?

