Video yükleniyor...
Video Yüklenemedi
Can we extend the power of world models beyond just online model-based learning? Absolutely! We believe the true potential of world models lies in enabling agents to reason at test time. Introducing DINO-WM: World Models on Pre-trained Visual Features for Zero-shot Planning.
168,129 görüntüleme • 1 yıl önce •via X (Twitter)
10 Yorum

Unlike previous works that couple world model learning with behavior learning, we train a dynamics-only model and infer actions only at test time. This allows zero-shot goal-reaching by reasoning through the dynamics—no expert demonstrations, no rewards, no online interactions.

DINO-WM consists of: 1️⃣An out-of-the-box DINOv2 model as the observation model. 2️⃣A causal ViT as the predictor. 3️⃣A decoder that is optional for visualization. DINO-WM plans entirely in latent space, without the need to reconstruct pixel images.

The object and spatial understanding priors of DINOv2 features enable robust scene understanding, essential for navigation and manipulation tasks. With this prior, DINO-WM outperforms state-of-the-art world models by 45% in downstream task performance on our hardest tasks.

Overall, DINO-WM takes a step toward bridging the gap between task-agnostic world modeling and reasoning and control, offering promising prospects for generic world models in real-world applications.

Huge thanks to all my collaborators who made this project possible @HengkaiPan , @ylecun , @LerrelPinto. We have open-sourced our code and data. For more details, checkout the paper and website: Website: arXiv:

great stuff Great team)

so basically what agents are doing in Web3 gaming, but irl

Cool!

how about inference time or frequence?

World models hold immense potential, but we must wield them with great care. Through a Foucauldian lens, I see these as powerful apparatuses that could reinforce existing power structures if not designed with rigorous ethical scrutiny.
