Загрузка видео...

Не удалось загрузить видео

На главную

Cosmos Policy Fine-Tuning Video Models for Visuomotor Control and Planning

10,355 просмотров • 8 месяцев назад •via X (Twitter)

Комментарии: 3

Фото профиля AK
AK8 месяцев назад

paper:

Фото профиля Moo Jin Kim
Moo Jin Kim8 месяцев назад

Thank you for sharing our work! 🙏

Фото профиля Nathan Wang
Nathan Wang8 месяцев назад

Video generation models aren’t just for making clips anymore—they’re becoming the robot brain itself. The part that really clicked for me in this NVIDIA/Stanford paper is the simplicity. Instead of bolting on complex control modules, they just treat robot actions as "latent frames" within the video sequence. Because of this, a single fine-tuned video model (Cosmos-Predict2) acts as the policy, the world simulator, and the value function all at once. It’s hitting 98.5% success on the LIBERO benchmark just by "dreaming" the right future frames and actions together. Makes you wonder—is the path to AGI just a really, really good video predictor?

Похожие видео