Загрузка видео...

Не удалось загрузить видео

На главную

Punchline: distill world models from simulation to enable fast, stable real-world robot adaptation. Simulation is nearly always wrong. But in Simulation Distillation, we ask a simple question: How do we perform simulation pretraining such that real-world adaptation becomes trivially easy? Let's take a closer look (1/n)

33,071 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 19

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Real-world RL is still pretty hard for long-horizon, contact-rich robotics. A few minutes of robot data is not enough to relearn rewards, values, representations, dynamics, and action selection end-to-end. And when we try, finetuning can destroy useful structure learned during pretraining. (2/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

The key idea in SimDist is simple: Move the hard RL problem into simulation, and transfer a high-coverage latent world model rather than only a policy. In simulation, we can train this model with privileged state, dense rewards, value functions, resets, perturbations, failures, and recoveries at scale. (3/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

At deployment, we can then freeze the parts that encode task structure — the encoder, reward model, and value function — and only adapt the latent dynamics model where the simulation imperfections lie. Performance is then naturally improved, just by performing short-horizon planning in this adapted model. So real-world adaptation becomes literally short-horizon supervised learning than long-horizon RL from scratch. (4/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

This is especially important for contact-rich tasks. The value function may already know that “peg aligned with hole” or “leg has stable foothold” is good, capturing the global structure of the task solution. What changes in the real world is often the details of local dynamics: friction, compliance, slip, contact, calibration. SimDist adapts exactly this part using real world data. (5/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Visualizing the value function is useful here, it transfers quite well and capture successes and failures as well as the orderings of which state is best. Importantly, this value function is trained with high-coverage, so only local repair of dynamics is needed - no OOD extrapolation errors! (6/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

One detail that matters during pre-training: Simulation data should not only contain successful expert trajectories. A planner asks counterfactual questions at test time: What if I push here? What if I slip? So the world model needs failures, perturbations, and recoveries too. This is where simulation is super useful. We can cheaply generate the messy off-policy experience that would be painful or unsafe to collect on real robots, then distill it into a model that supports real-world planning and rapid adaptation. (7/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Across precise manipulation and quadruped locomotion, SimDist adapts from just 15-30 min of real-world data by planning with transferred reward/value structure and finetuning only dynamics. The adaptation is stable because it is supervised learning, rather than online RL over the full policy/value/reward stack. (8/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

The perspective I find useful: Simulation gives broad structure and coverage, real-world data fixes the local details. SimDist reduces adaptation to just what needs to change, while avoiding many of the extrapolation and instability issues that appear in offline-to-online RL finetuning. (9/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Big kudos to Jacob Levy and @ty_westenbroek for leading this, along with Kevin Huang, Fernando Palafox, @patrickhyin , Dong-Ki Kim, Shayegan Omidshafiei, and David Fridovich-Keil. They also made a fantastic website to help understand the work, where you can play with model behavior and performance: Website: Paper: See you at RSS 2026 to talk more about this work! (10/n)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Special note to highlight that @ty_westenbroek is on the job market for research scientist roles! Hire him, he does fantastic work :)

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Sim2real 🤝 World Models 🤝 RL, we have all the buzzwords! 🎉🎉🎉🥳

Фото профиля Vikash Kumar
Vikash Kumar4 месяцев назад

Do we need to know the task distribution, or can it be trained agnostic to the distribution ahead of times?

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Currently with a known task distribution, since it needs to get the VFs in simulation. But would be fun to get rid of that requirement :)

Фото профиля Dr. Richard
Dr. Richard4 месяцев назад

I love how this tackles the sim-to-real bottleneck. Latent world models distilled from simulation are exactly what robust industrial automation needs. Great work 👏

Фото профиля Chenhao Li
Chenhao Li4 месяцев назад

A very similar idea as Robotic World Model but with images! Cool work!

Фото профиля Abhishek Gupta
Abhishek Gupta4 месяцев назад

Thanks! :) I think the architecture itself is pretty close, but I want to highlight that the primary focus of SimDist is around real-world adaptation post deployment. So it's about finetuning the world model itself using small amounts of real-world data after pretraining a bunch in simulation. As opposed to a better policy optimization procedure, like robotic world models - SimDist is focused on fast test-time finetuning in the real world. So I think they can be nicely complementary :)

Фото профиля Chenhao Li
Chenhao Li4 месяцев назад

I believe in Robotic World Model, the model is designed to be finetuned during online adaptations, then the policy is further finetuned within this model. There’s an offline version of it, Uncertain Aware Robotic World Model. You might have been referring to that one. 🤔

Фото профиля Mike Carrieri
Mike Carrieri4 месяцев назад

#266 9/30/2019

Фото профиля Saïd Aitmbarek
Saïd Aitmbarek4 месяцев назад

based

Похожие видео