Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Cosmos Policy Fine-Tuning Video Models for Visuomotor Control and Planning

10,355 Aufrufe • vor 8 Monaten •via X (Twitter)

3 Kommentare

Profilbild von AK
AKvor 8 Monaten

paper:

Profilbild von Moo Jin Kim
Moo Jin Kimvor 8 Monaten

Thank you for sharing our work! 🙏

Profilbild von Nathan Wang
Nathan Wangvor 8 Monaten

Video generation models aren’t just for making clips anymore—they’re becoming the robot brain itself. The part that really clicked for me in this NVIDIA/Stanford paper is the simplicity. Instead of bolting on complex control modules, they just treat robot actions as "latent frames" within the video sequence. Because of this, a single fine-tuned video model (Cosmos-Predict2) acts as the policy, the world simulator, and the value function all at once. It’s hitting 98.5% success on the LIBERO benchmark just by "dreaming" the right future frames and actions together. Makes you wonder—is the path to AGI just a really, really good video predictor?

Ähnliche Videos