Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Generalist robots need a generalist evaluator. But how do you test safety without breaking things? 💥 🌎 Introducing our new work from Google DeepMind: Evaluating Gemini Robotics Policies in a Veo World Simulator 🧵👇

242,936 görüntüleme • 9 ay önce •via X (Twitter)

31 Yorum

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

The challenge: hardware evaluation is slow and often dangerous for safety probing. Traditional physics sims are great but struggle with visual realism, non-rigid objects, and the immense manual effort required for accuracy. Can video models serve as a generalist simulator?

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

💡 We introduce a system that utilizes @GoogleDeepMind’s Veo to evaluate policies on: ✅ Nominal performance on real-world tasks ✅ Out-of-Distribution (OOD) generalization ✅ Semantic Safety All without a single physics engine.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

🛠️ We fine-tuned Veo for action conditioning & multi-view generation. This was achieved using a large-scale robotics dataset containing a wide range of different scenes and tasks.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

We use the action-conditioned world model to roll out a given policy in an initial scene. We can score the policy rollout to predict real-world success or failure.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

🤯 A surprise: we found that fine-tuning the video model didn't break its "common sense." We can edit scenes to include novel objects or distractors, and still simulate the physics and interactions, effectively preserving the base model's world knowledge.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

📊 Validation: We ran 1600+ real-world evaluations across 8 Gemini Robotics policy checkpoints and 5 tasks. Result: the Veo (Robotics) model accurately predicts the relative performance and rankings of different policies in both nominal and OOD conditions.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

🛡️ Predictive Red Teaming: We used Veo to probe semantic safety. We can simulate dangerous "long tail" scenarios without breaking things in real life. 🚫 We replicated scenarios using props and found that predicted behaviors were real.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

This work represents a step toward scalable, safe, and generalist evaluation for the next generation of robot policies using world models. 📜Read the full report here: 💻Website:

Saleh Aldwais profil fotoğrafı
Saleh Aldwais9 ay önce

@GoogleDeepMind Thanks for sharing

Jack Monas profil fotoğrafı
Jack Monas9 ay önce

@GoogleDeepMind Really cool work Ani! Great to see more results on WM for evals. Curious how much of an improvement a Veo3 base model would provide for physical alignment compared to Veo2.

Anirudha Majumdar profil fotoğrafı
Anirudha Majumdar9 ay önce

@GoogleDeepMind Thanks so much, @JackMonas! It’s been awesome seeing your work at 1X on these topics

satvik profil fotoğrafı
satvik9 ay önce

@GoogleDeepMind WOW!!!!!!!!! really wish you guys could open source this

Serena Yan profil fotoğrafı
Serena Yan9 ay önce

@GoogleDeepMind How long does each eval take, and running on what gpu?

anon profil fotoğrafı
anon9 ay önce

@GoogleDeepMind Wat Guess veo is a physics model ¯\_(ツ)_/¯

Satpal Singh Rathore profil fotoğrafı
Satpal Singh Rathore9 ay önce

@GoogleDeepMind Isn't 80 scene instruction bit low to validate this?

Today in AI profil fotoğrafı
Today in AI9 ay önce

@GoogleDeepMind Generalist policy evaluation requires Veo fidelity to hold across an unbounded state-action space; the sim-to-real gap magnifies non-linearly with policy generality.

Buck.Sol profil fotoğrafı
Buck.Sol9 ay önce

@GoogleDeepMind simulation first

Leon profil fotoğrafı
Leon9 ay önce

@GoogleDeepMind sim first then field later makes a lotta sense here

Lijuan profil fotoğrafı
Lijuan2 ay önce

@GoogleDeepMind This is exactly the missing piece for safe general robotics development

Ali.BTC profil fotoğrafı
Ali.BTC9 ay önce

@GoogleDeepMind Robots dreaming before they wake

Today in AI profil fotoğrafı
Today in AI9 ay önce

@GoogleDeepMind Sim-to-real fidelity remains the critical safety bottleneck. Veo must prove it captures edge-case physics and catastrophic failure modes, not just nominal 0.88 correlation on ALOHA 2 tasks.

Jeycosmos ⚔️ profil fotoğrafı
Jeycosmos ⚔️9 ay önce

@GoogleDeepMind this looks like a smart way to test safely without risking real damage

Satpal Singh Rathore profil fotoğrafı
Satpal Singh Rathore9 ay önce

@GoogleDeepMind Useful to identify edge case failures, but then how do you know these evals match real world eval results even for edge cases.

Von.ETH profil fotoğrafı
Von.ETH9 ay önce

@GoogleDeepMind wild to think it works

Lukas profil fotoğrafı
Lukas9 ay önce

@GoogleDeepMind cool work! could this be used beyond evaluation to also fine tune policies?

Amy - AI Girl profil fotoğrafı
Amy - AI Girl9 ay önce

@GoogleDeepMind Exciting approach to testing safety in robotics! Looking forward to seeing the results.

zdapricorn.eth profil fotoğrafı
zdapricorn.eth9 ay önce

@GoogleDeepMind feels like sci‑fi becoming baseline

Naz Louis profil fotoğrafı
Naz Louis9 ay önce

@GoogleDeepMind Made this using Gemini 1.5 ER, can’t wait for the VLA model to go public so I can implement so more complex task.

Himanshu Kumar profil fotoğrafı
Himanshu Kumar9 ay önce

@GoogleDeepMind Anirudha, it's a valid point: evaluating robot safety is a real challenge. You'll find the Veo World Simulator work insightful for that. Useful reference for generalist:

AI PlanetX profil fotoğrafı
AI PlanetX9 ay önce

@GoogleDeepMind Exciting work! Testing safety in a Veo World Simulator is innovative. Can't wait to see the results!

Ankita Kale profil fotoğrafı
Ankita Kale9 ay önce

@GoogleDeepMind So easy to train for situations that are risky and hard to create.

Benzer Videolar