Video wird geladen...
Video konnte nicht geladen werden
🌎World models can predict, but controlling real robots from imagination sees a long-standing failure due to hallucination. 🧠Introducing Uncertainty-Aware RWM: a black-box, end-to-end neural dynamics model with long-horizon uncertainty propagation. 🎯
60,168 Aufrufe • vor 8 Monaten •via X (Twitter)
16 Kommentare

❗️World models hallucinate under distribution shift: in low-data regions, small errors compound over long autoregressive rollouts. 🧠Policy optimization then exploits these hallucinations, achieving high reward in imagination while failing catastrophically in the real world.

🌎Uncertainty-Aware Robotic World Model (RWM-U) extends RWM by modeling not only what will happen, but how reliable those predictions are. 🎯We augment autoregressive world models with ensemble-based uncertainty estimation, explicitly capturing epistemic uncertainty from limited or biased offline data.

⭐️Each ensemble member in RWM-U predicts a Gaussian distribution over the next observation. 🎯The predicted variance captures aleatoric uncertainty, while disagreement across ensemble means estimates epistemic uncertainty from limited or biased offline data.

⭐️In long autoregressive rollouts, RWM-U’s epistemic uncertainty closely tracks true prediction error, providing a reliable signal of when the model becomes untrustworthy.

We bring MOPO to long-horizon world models, and to PPO. ✅MOPO-PPO trains policies entirely in imagination, penalizing uncertain transitions to avoid hallucination. No real environment interaction ⏩fast, fully offline learning.

💾Starting from pure offline data, with no online environment interaction (not even a simulator), we train policies that are directly deployable on real hardware.

⭐️Though completely offline, RWM-U and MOPO-PPO achieve comparable performance as state-of-the-art simulator-based online policies.

🌎A key advantage of model-based RL is that the world model can be trained directly on real-world data. 🔝By incorporating real data into model learning, we reduce the sim-to-real gap and achieve final policy performance that surpasses sim-to-real baselines.

🧠Policy conservativeness is controlled by the uncertainty penalty weight during imagination. 🛑When offline data has limited task relevance, increasing the penalty yields more conservative and safer policies.

🌎Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots 👥Chenhao Li @breadli428, Andreas Krause, Marco Hutter @leggedrobotics 🎯Webpage: 📄Paper:

Nice

love seeing this. great work addressing robot hallucination with better models.

The black-box neural dynamics approach here intrigues me. How does its uncertainty propagation compare to traditional control systems in real-world robotics applications?

Dm

how does uncertainty propagation impact real-world robot control stability?

so they predict hallucinations now? sounds like just another overhyped ai buzzword waiting to fail in real applications.
