Loading video...

Video Failed to Load

Go Home

Introducing World Tracing Generative pixel-aligned geometry, beyond the visible. Faithful to your pixels. Complete in 3D. One image in — objects, scenes, even dynamic worlds emerge in full geometry, every point traced back to the pixel it came from. 🌐 🧵 [1/6]

106,299 views • 3 months ago •via X (Twitter)

18 Comments

Hao Zhang's profile picture
Hao Zhang3 months ago

Single-image to 3D forced a choice: depth is faithful but stops at the surface, or generation is complete but drifts from your pixels. World Tracing dissolves the trade-off, each pixel ray carries an ordered stack of 3D points, from the visible surface to the geometry hidden behind it. One tensor. One model. Reconstruction and generation, unified.

Hao Zhang's profile picture
Hao Zhang3 months ago

Paper: Page: Code: Live Demo:

Hao Zhang's profile picture
Hao Zhang3 months ago

Completion comes at no cost to fidelity: WT's first layer alone matches SOTA monocular depth predictors on visible-surface accuracy, while on faithful generation, complete geometry aligned to the input image, it outperforms canonical-frame image-to-3D and video-to-4D generators.

Hao Zhang's profile picture
Hao Zhang3 months ago

And because every 3D point remembers its pixel, WT becomes a universal geometry interface: 3D scene editing, geometry-guided video synthesis, training-free textured meshes.

Hao Zhang's profile picture
Hao Zhang3 months ago

Big thanks to all collaborators and people support this project: @theworldlabs @gengshanY @BenMildenhall @chlassner @jcjohnss @KeunhongP @_mbanani @DrTunnels @Hawaii271828 @PaulZhang @cuijiaxingfb @BDuisterhof @zixuan_huang @drfeifei

Hao Zhang's profile picture
Hao Zhang3 months ago

Thanks for trying the model! It seems the current released model doesn't work very well on human face details. I think the incoimg larger model will perform better with those details.

Ryan Schmidt's profile picture
Ryan Schmidt3 months ago

this is cool, nice work

Ian Curtis's profile picture
Ian Curtis3 months ago

Excited to try this out!

SomacoSF's profile picture
SomacoSF3 months ago

Imagine throwing /flying a ball/drone into a space o it just flies through SUPER fast and snaps bajillions of photos that then get sent back to a this bot and it gives you full 3d of the environ asap (tiny race drones that fly >100KMHH style recon sortie

AUTONOMOUS's profile picture
AUTONOMOUS3 months ago

Excited to see this hopefully being used in data training moving forward

Ziyang Xie's profile picture
Ziyang Xie3 months ago

Amazing work!

平地有樹曰林's profile picture
平地有樹曰林3 months ago

wow

Shikun Liu's profile picture
Shikun Liu3 months ago

Beautiful work and beautiful project page!

SomacoSF's profile picture
SomacoSF3 months ago

Acrually you could have this fly in a drove and map out trail conditions and share it on an app that allows bikers hikers to load the terrain into thei system and have it know all the details of a trail in splats

Yo.Town 3D's profile picture
Yo.Town 3D3 months ago

This facilitates the creation of more 3D models. In cases where 3D printing is required, we may divide the models like,

Yunfan's profile picture
Yunfan3 months ago

为什么要这么做呢? 从训练角度来看 人的3D视觉很有可能是因为两路输入信号(左右眼)的问题 但是我看你们学术界一直都是在玩类似单眼输入的映射

SomacoSF's profile picture
SomacoSF3 months ago

you could also fly a drone ahead of say some smart EV/E-Bike along a trail and have it basically trace/lidar the terrain and send the exact model of the terrain to the chasing ev and it maps the suspension exactly with the terrain allowing for accelerationmaxxing BAJA style

Hao Zhang's profile picture
Hao Zhang3 months ago

Welcome to try our model and share the failure cases! We will continue updating the more powerful checkpoints and models!

Related Videos