Video yükleniyor...
Video Yüklenemedi
Generalist robots need a generalist evaluator. But how do you test safety without breaking things? 💥 🌎 Introducing our new work from Google DeepMind: Evaluating Gemini Robotics Policies in a Veo World Simulator 🧵👇
242,936 görüntüleme • 9 ay önce •via X (Twitter)
31 Yorum

The challenge: hardware evaluation is slow and often dangerous for safety probing. Traditional physics sims are great but struggle with visual realism, non-rigid objects, and the immense manual effort required for accuracy. Can video models serve as a generalist simulator?

💡 We introduce a system that utilizes @GoogleDeepMind’s Veo to evaluate policies on: ✅ Nominal performance on real-world tasks ✅ Out-of-Distribution (OOD) generalization ✅ Semantic Safety All without a single physics engine.

🛠️ We fine-tuned Veo for action conditioning & multi-view generation. This was achieved using a large-scale robotics dataset containing a wide range of different scenes and tasks.

We use the action-conditioned world model to roll out a given policy in an initial scene. We can score the policy rollout to predict real-world success or failure.

🤯 A surprise: we found that fine-tuning the video model didn't break its "common sense." We can edit scenes to include novel objects or distractors, and still simulate the physics and interactions, effectively preserving the base model's world knowledge.

📊 Validation: We ran 1600+ real-world evaluations across 8 Gemini Robotics policy checkpoints and 5 tasks. Result: the Veo (Robotics) model accurately predicts the relative performance and rankings of different policies in both nominal and OOD conditions.

🛡️ Predictive Red Teaming: We used Veo to probe semantic safety. We can simulate dangerous "long tail" scenarios without breaking things in real life. 🚫 We replicated scenarios using props and found that predicted behaviors were real.

This work represents a step toward scalable, safe, and generalist evaluation for the next generation of robot policies using world models. 📜Read the full report here: 💻Website:

@GoogleDeepMind Thanks for sharing

@GoogleDeepMind Really cool work Ani! Great to see more results on WM for evals. Curious how much of an improvement a Veo3 base model would provide for physical alignment compared to Veo2.

@GoogleDeepMind Thanks so much, @JackMonas! It’s been awesome seeing your work at 1X on these topics

@GoogleDeepMind WOW!!!!!!!!! really wish you guys could open source this

@GoogleDeepMind How long does each eval take, and running on what gpu?

@GoogleDeepMind Wat Guess veo is a physics model ¯\_(ツ)_/¯

@GoogleDeepMind Isn't 80 scene instruction bit low to validate this?

@GoogleDeepMind Generalist policy evaluation requires Veo fidelity to hold across an unbounded state-action space; the sim-to-real gap magnifies non-linearly with policy generality.

@GoogleDeepMind simulation first

@GoogleDeepMind sim first then field later makes a lotta sense here

@GoogleDeepMind This is exactly the missing piece for safe general robotics development

@GoogleDeepMind Robots dreaming before they wake

@GoogleDeepMind Sim-to-real fidelity remains the critical safety bottleneck. Veo must prove it captures edge-case physics and catastrophic failure modes, not just nominal 0.88 correlation on ALOHA 2 tasks.

@GoogleDeepMind this looks like a smart way to test safely without risking real damage

@GoogleDeepMind Useful to identify edge case failures, but then how do you know these evals match real world eval results even for edge cases.

@GoogleDeepMind wild to think it works

@GoogleDeepMind cool work! could this be used beyond evaluation to also fine tune policies?

@GoogleDeepMind Exciting approach to testing safety in robotics! Looking forward to seeing the results.

@GoogleDeepMind feels like sci‑fi becoming baseline

@GoogleDeepMind Made this using Gemini 1.5 ER, can’t wait for the VLA model to go public so I can implement so more complex task.

@GoogleDeepMind Anirudha, it's a valid point: evaluating robot safety is a real challenge. You'll find the Veo World Simulator work insightful for that. Useful reference for generalist:

@GoogleDeepMind Exciting work! Testing safety in a Veo World Simulator is innovative. Can't wait to see the results!

@GoogleDeepMind So easy to train for situations that are risky and hard to create.
