正在加载视频...

视频加载失败

🌎World models can predict, but controlling real robots from imagination sees a long-standing failure due to hallucination. 🧠Introducing Uncertainty-Aware RWM: a black-box, end-to-end neural dynamics model with long-horizon uncertainty propagation. 🎯

60,168 次观看 • 8 个月前 •via X (Twitter)

16 条评论

Chenhao Li 的头像
Chenhao Li8 个月前

❗️World models hallucinate under distribution shift: in low-data regions, small errors compound over long autoregressive rollouts. 🧠Policy optimization then exploits these hallucinations, achieving high reward in imagination while failing catastrophically in the real world.

Chenhao Li 的头像
Chenhao Li8 个月前

🌎Uncertainty-Aware Robotic World Model (RWM-U) extends RWM by modeling not only what will happen, but how reliable those predictions are. 🎯We augment autoregressive world models with ensemble-based uncertainty estimation, explicitly capturing epistemic uncertainty from limited or biased offline data.

Chenhao Li 的头像
Chenhao Li8 个月前

⭐️Each ensemble member in RWM-U predicts a Gaussian distribution over the next observation. 🎯The predicted variance captures aleatoric uncertainty, while disagreement across ensemble means estimates epistemic uncertainty from limited or biased offline data.

Chenhao Li 的头像
Chenhao Li8 个月前

⭐️In long autoregressive rollouts, RWM-U’s epistemic uncertainty closely tracks true prediction error, providing a reliable signal of when the model becomes untrustworthy.

Chenhao Li 的头像
Chenhao Li8 个月前

We bring MOPO to long-horizon world models, and to PPO. ✅MOPO-PPO trains policies entirely in imagination, penalizing uncertain transitions to avoid hallucination. No real environment interaction ⏩fast, fully offline learning.

Chenhao Li 的头像
Chenhao Li8 个月前

💾Starting from pure offline data, with no online environment interaction (not even a simulator), we train policies that are directly deployable on real hardware.

Chenhao Li 的头像
Chenhao Li8 个月前

⭐️Though completely offline, RWM-U and MOPO-PPO achieve comparable performance as state-of-the-art simulator-based online policies.

Chenhao Li 的头像
Chenhao Li8 个月前

🌎A key advantage of model-based RL is that the world model can be trained directly on real-world data. 🔝By incorporating real data into model learning, we reduce the sim-to-real gap and achieve final policy performance that surpasses sim-to-real baselines.

Chenhao Li 的头像
Chenhao Li8 个月前

🧠Policy conservativeness is controlled by the uncertainty penalty weight during imagination. 🛑When offline data has limited task relevance, increasing the penalty yields more conservative and safer policies.

Chenhao Li 的头像
Chenhao Li8 个月前

🌎Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots 👥Chenhao Li @breadli428, Andreas Krause, Marco Hutter @leggedrobotics 🎯Webpage: 📄Paper:

ζ Pedram ζ 的头像
ζ Pedram ζ8 个月前

Nice

Sohom Mukherjee 的头像
Sohom Mukherjee8 个月前

love seeing this. great work addressing robot hallucination with better models.

Lukas Die Kunst 的头像
Lukas Die Kunst8 个月前

The black-box neural dynamics approach here intrigues me. How does its uncertainty propagation compare to traditional control systems in real-world robotics applications?

Damian 的头像
Damian8 个月前

Dm

Rahul Raghav 的头像
Rahul Raghav8 个月前

how does uncertainty propagation impact real-world robot control stability?

Eddie 的头像
Eddie8 个月前

so they predict hallucinations now? sounds like just another overhyped ai buzzword waiting to fail in real applications.

相关视频

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 次观看 • 2 个月前