Loading video...
Video Failed to Load
Introducing 👀Stereo4D👀 A method for mining 4D from internet stereo videos. It enables large-scale, high-quality, dynamic, *metric* 3D reconstructions, with camera poses and long-term 3D motion trajectories. We used Stereo4D to make a dataset of over 100k real-world 4D scenes.
94,683 views • 1 year ago •via X (Twitter)
16 Comments

This type of data is ideal for learning the structure and dynamics of the real world. We gave this a shot — extending DUSt3R to model 3D motion, and training on our dataset. Given a pair of frames, our model predicts a 3D point cloud, and corresponding 3D motion trajectories.

See more scenes & details of how it works on our website: Paper: Thanks to the great team! Richard Tucker, @zhengqi_li, David Fouhey, @Jimantha, @holynski_ Please stay tuned for updates on data & code.

Congrats, will be super useful for the 4D reconstruction/generation community!!

Very amazing work! Wondering how to predict the 3D point trajectory of a video after pairwise prediction of DynaDust3r?

Thanks Chen! We've tried using DynaDust3r to predict motion at any time between two frames. We haven't tried to extend it to take all frames of a video, but could be a cool future direction.

Excellent paper @jin_linyi ! Exciting potential for models leveraging multi-modal data 🦾

Very interesting work!

This should be Good for 3D tracking . @jin_linyi . both 3D camera tracking and 3D mesh tracking.

great work! when will it be released? can't wait to use it.

🤩 wow 🤩

The world simulation and the singularity is near @PeterDiamandis @salimismail

Me patiently waiting for code

@zhengqi_li Time is Light !

Cool!

I'm blown away by the scale of your dataset! How do you plan to make it available for others to use?

Really cool to see VR180 videos have this impact -- this'll be huge for the computer vision community!
