正在加载视频...
视频加载失败
What would a World Model look like if we start from a real embodied agent acting in the real world? It has to have: 1) A real, physically grounded and complex action space—not just abstract control signals. 2) Diverse, real-life scenarios and activities. Or in short: It has to... show more
10 条评论

Let’s see some pixels! We mainly showcase three types of results - atom actions, long video rollouts and counterfactuals for planning.

For more atom actions:

For more long video rolling out examples

For counterfactual planning:

Here’s the project page with more results, failures and discussions we have, check these out!

@ylecun @JitendraMalikCV @trevordarrell @_amirbar @dans_t123 Integrating the complex action space with grounded physicality could yield more robust predictions compared to existing abstract models.

Exactly — from a first-person human perspective, even something as simple as walking causes the first-person video to shake along with the person’s gait and movement. So if the model can learn the relationship between this motion and the video, it could potentially be more robust to that shaking when understanding the goal.

@ylecun @JitendraMalikCV @trevordarrell @_amirbar @dans_t123 Integrating embodiment adds significant modeling subtleties. Promising direction.

this is very interesting and timely @GnosisYu did some work on egocentric camera control recently at ICLR -- EgoSim wondering what comparisons may look like as well as what architectural changes were needed for precise camera control. @ylecun, @trevordarrell, @JitendraMalikCV

@ylecun @JitendraMalikCV @trevordarrell @_amirbar @dans_t123 @GnosisYu Thank you for pointing this out! We will cite this in the next version. Yeah camera pose is very important. I think human’s complex body motion make this problem even harder.
