Video yükleniyor...
Video Yüklenemedi
World models have a causality problem. Realistic videos are not enough. A world model should predict the future caused by an action, not just a plausible future. We find that many latent-action world models generate convincing videos while barely responding to the supplied action. The root cause lies in... show more
15,017 görüntüleme • 2 ay önce •via X (Twitter)
9 Yorum

Technical blog: Paper: Project page & videos: Code: Models:

This is so good!! Thanks for sharing.

turns out they just hallucinate physics

Right diagnosis, and it goes deeper: the benchmarks reward the failure. FVD scores plausibility, so a model that ignores the action but renders beautifully still wins. Until action-following is the headline metric, LAMs keep optimizing for the wrong thing.

The reality humans live in is not a statistical average but an exact measurement.

Does better causality here directly translate to better robot performance, or is there still another bottleneck after that?

Does this work if the causality has a time delay?

My take: causality is the bottleneck, not realism. CD-LAM fixing latent confounding pre-training and cutting 50k steps to 3k is exactly the leap embodied AI needs. Solid work.

no
