Video yükleniyor...
Video Yüklenemedi
Introducing PointZero—a 3D world model pre-trained without robots. Dexterous manipulation requires understanding diverse 3D dynamics, but current approaches rely on expensive robot data. So how can we pre-train a 3D dynamics model without it? PointZero introduces a simple idea: learning to complete 3D point tracks yields transferable 3D dynamics!... show more
27,011 görüntüleme • 3 gün önce •via X (Twitter)
14 Yorum

Our approach is simple: given an RGB-D observation and only 1-3 tracks, PointZero denoises point tracks for all observed points. Sparse tracks constrain possible outcomes, but leave properties such as stiffness and joint structure ambiguous. 2/6

Trained on just 2.9M frames of diverse synthetic dynamics, PointZero produces plausible zero-shot predictions of real-world dynamics. We test it on deformable, articulated, and rigid objects. 3/6

We post-train PointZero for two downstream applications. First, action-conditioned 3D dynamics: instead of point tracks, we condition on the robot’s end-effector state. PointZero outperforms the task specialists we evaluate. 4/6

Next, we turn PointZero into an MM-DiT to predict robot actions through imitation learning. It matches or outperforms existing behavior cloning methods. We also compare pre-trained PointZero with training from scratch: pre-training improves performance in both downstream applications. 5/6

Thanks to awesome co-authors @kaiwynd @Adamjhung @bowenwen_me @BirchfieldStan @YunzhuLiYZ @RamananDeva @jeff_ichnowski The paper is out now; the code and dataset will follow shortly. Website: arXiv: 6/6

I like this! @bardienus

Thanks prof @mangahomanga !

bro is on rampage

Lol release week 😄

Hahah do you have more ? 👀

Just this one:

😂😂😂 my man

wild scaling

One to three tracks can leave several motions plausible, especially with hidden joints or stiffness. Preserving that uncertainty seems important if a robot is going to choose its next action from the prediction.
