Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Cosmos Policy Fine-Tuning Video Models for Visuomotor Control and Planning

10,355 görüntüleme • 8 ay önce •via X (Twitter)

3 Yorum

AK profil fotoğrafı
AK8 ay önce

paper:

Moo Jin Kim profil fotoğrafı
Moo Jin Kim8 ay önce

Thank you for sharing our work! 🙏

Nathan Wang profil fotoğrafı
Nathan Wang8 ay önce

Video generation models aren’t just for making clips anymore—they’re becoming the robot brain itself. The part that really clicked for me in this NVIDIA/Stanford paper is the simplicity. Instead of bolting on complex control modules, they just treat robot actions as "latent frames" within the video sequence. Because of this, a single fine-tuned video model (Cosmos-Predict2) acts as the policy, the world simulator, and the value function all at once. It’s hitting 98.5% success on the LIBERO benchmark just by "dreaming" the right future frames and actions together. Makes you wonder—is the path to AGI just a really, really good video predictor?

Benzer Videolar