Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🌎World models can predict, but controlling real robots from imagination sees a long-standing failure due to hallucination. 🧠Introducing Uncertainty-Aware RWM: a black-box, end-to-end neural dynamics model with long-horizon uncertainty propagation. 🎯

60,168 Aufrufe • vor 8 Monaten •via X (Twitter)

16 Kommentare

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

❗️World models hallucinate under distribution shift: in low-data regions, small errors compound over long autoregressive rollouts. 🧠Policy optimization then exploits these hallucinations, achieving high reward in imagination while failing catastrophically in the real world.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

🌎Uncertainty-Aware Robotic World Model (RWM-U) extends RWM by modeling not only what will happen, but how reliable those predictions are. 🎯We augment autoregressive world models with ensemble-based uncertainty estimation, explicitly capturing epistemic uncertainty from limited or biased offline data.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

⭐️Each ensemble member in RWM-U predicts a Gaussian distribution over the next observation. 🎯The predicted variance captures aleatoric uncertainty, while disagreement across ensemble means estimates epistemic uncertainty from limited or biased offline data.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

⭐️In long autoregressive rollouts, RWM-U’s epistemic uncertainty closely tracks true prediction error, providing a reliable signal of when the model becomes untrustworthy.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

We bring MOPO to long-horizon world models, and to PPO. ✅MOPO-PPO trains policies entirely in imagination, penalizing uncertain transitions to avoid hallucination. No real environment interaction ⏩fast, fully offline learning.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

💾Starting from pure offline data, with no online environment interaction (not even a simulator), we train policies that are directly deployable on real hardware.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

⭐️Though completely offline, RWM-U and MOPO-PPO achieve comparable performance as state-of-the-art simulator-based online policies.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

🌎A key advantage of model-based RL is that the world model can be trained directly on real-world data. 🔝By incorporating real data into model learning, we reduce the sim-to-real gap and achieve final policy performance that surpasses sim-to-real baselines.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

🧠Policy conservativeness is controlled by the uncertainty penalty weight during imagination. 🛑When offline data has limited task relevance, increasing the penalty yields more conservative and safer policies.

Profilbild von Chenhao Li
Chenhao Livor 8 Monaten

🌎Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots 👥Chenhao Li @breadli428, Andreas Krause, Marco Hutter @leggedrobotics 🎯Webpage: 📄Paper:

Profilbild von ζ Pedram ζ
ζ Pedram ζvor 8 Monaten

Nice

Profilbild von Sohom Mukherjee
Sohom Mukherjeevor 8 Monaten

love seeing this. great work addressing robot hallucination with better models.

Profilbild von Lukas Die Kunst
Lukas Die Kunstvor 8 Monaten

The black-box neural dynamics approach here intrigues me. How does its uncertainty propagation compare to traditional control systems in real-world robotics applications?

Profilbild von Damian
Damianvor 8 Monaten

Dm

Profilbild von Rahul Raghav
Rahul Raghavvor 8 Monaten

how does uncertainty propagation impact real-world robot control stability?

Profilbild von Eddie
Eddievor 8 Monaten

so they predict hallucinations now? sounds like just another overhyped ai buzzword waiting to fail in real applications.

Ähnliche Videos

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 Aufrufe • vor 2 Monaten