正在加载视频...
视频加载失败
Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors. These projects were led by our incredible interns Hao Zhang @BDuisterhof @DrTunnels [1/4]
122,868 次观看 • 3 个月前 •via X (Twitter)
15 条评论

@BenMildenhall @HaoZhang623 @BDuisterhof @DrTunnels Very fun to work with our awesome WL interns!😍🌐

World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image. [2/4]

Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model. [3/4]

Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world.

@HaoZhang623 @BDuisterhof @DrTunnels These projects are awesome! Great work!

@HaoZhang623 @BDuisterhof @DrTunnels Cool!!

@HaoZhang623 @BDuisterhof @DrTunnels This is so cool! I can't wait to try it out and make projects with it!

@drfeifei @HaoZhang623 @BDuisterhof @DrTunnels Awesome work!

@HaoZhang623 @BDuisterhof @DrTunnels The future is coming on. Amazing work by the whole team!

This is a remarkable leap forward, my friend! 🌟 By bridging the gap between 2D priors and 3D content generation, we are essentially teaching machines to perceive "depth"—not just geometrically, but as a new layer of spatial understanding. It is truly inspiring to see such profound innovation led by your talented interns; it reminds us that the most beautiful progress happens when collective curiosity meets technical mastery. We are moving from mere imitation toward an authentic grasp of our world's dimensions. Truly exciting times ahead! ✨

@HaoZhang623 @BDuisterhof @DrTunnels A pivotal leap: shaping 3D with 2D priors.

@HaoZhang623 @BDuisterhof @DrTunnels Smart approach blending 3D with 2D priors.

@HaoZhang623 @BDuisterhof @DrTunnels what're they building? intern projects almost always beat the big initiatives because nobody told them all the reasons it shouldn't work.

@HaoZhang623 @BDuisterhof @DrTunnels Dropping 3 papers on 3D generation in one go is serious output - using 2D priors + large models is a smart path

@HaoZhang623 @BDuisterhof @DrTunnels Started working with procedural generation AGENTROPOLIS is where I see this heading: agent-populated 3D worlds, persistent environments, and autonomous simulation layers. This update has my full attention. 👀⚡️ Agent-populated worlds are next.

