Loading video...
Video Failed to Load
Introducing World Tracing Generative pixel-aligned geometry, beyond the visible. Faithful to your pixels. Complete in 3D. One image in — objects, scenes, even dynamic worlds emerge in full geometry, every point traced back to the pixel it came from. 🌐 🧵 [1/6]
106,299 views • 3 months ago •via X (Twitter)
18 Comments

Single-image to 3D forced a choice: depth is faithful but stops at the surface, or generation is complete but drifts from your pixels. World Tracing dissolves the trade-off, each pixel ray carries an ordered stack of 3D points, from the visible surface to the geometry hidden behind it. One tensor. One model. Reconstruction and generation, unified.

Paper: Page: Code: Live Demo:

Completion comes at no cost to fidelity: WT's first layer alone matches SOTA monocular depth predictors on visible-surface accuracy, while on faithful generation, complete geometry aligned to the input image, it outperforms canonical-frame image-to-3D and video-to-4D generators.

And because every 3D point remembers its pixel, WT becomes a universal geometry interface: 3D scene editing, geometry-guided video synthesis, training-free textured meshes.

Big thanks to all collaborators and people support this project: @theworldlabs @gengshanY @BenMildenhall @chlassner @jcjohnss @KeunhongP @_mbanani @DrTunnels @Hawaii271828 @PaulZhang @cuijiaxingfb @BDuisterhof @zixuan_huang @drfeifei

Thanks for trying the model! It seems the current released model doesn't work very well on human face details. I think the incoimg larger model will perform better with those details.

this is cool, nice work

Excited to try this out!

Imagine throwing /flying a ball/drone into a space o it just flies through SUPER fast and snaps bajillions of photos that then get sent back to a this bot and it gives you full 3d of the environ asap (tiny race drones that fly >100KMHH style recon sortie

Excited to see this hopefully being used in data training moving forward

Amazing work!

wow

Beautiful work and beautiful project page!

Acrually you could have this fly in a drove and map out trail conditions and share it on an app that allows bikers hikers to load the terrain into thei system and have it know all the details of a trail in splats

This facilitates the creation of more 3D models. In cases where 3D printing is required, we may divide the models like,

为什么要这么做呢? 从训练角度来看 人的3D视觉很有可能是因为两路输入信号(左右眼)的问题 但是我看你们学术界一直都是在玩类似单眼输入的映射

you could also fly a drone ahead of say some smart EV/E-Bike along a trail and have it basically trace/lidar the terrain and send the exact model of the terrain to the chasing ev and it maps the suspension exactly with the terrain allowing for accelerationmaxxing BAJA style

Welcome to try our model and share the failure cases! We will continue updating the more powerful checkpoints and models!

