正在加载视频...

视频加载失败

Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors. These projects were led by our incredible interns Hao Zhang @BDuisterhof @DrTunnels [1/4]

122,868 次观看 • 3 个月前 •via X (Twitter)

15 条评论

Fei-Fei Li 的头像
Fei-Fei Li3 个月前

@BenMildenhall @HaoZhang623 @BDuisterhof @DrTunnels Very fun to work with our awesome WL interns!😍🌐

World Labs 的头像
World Labs3 个月前

World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image. [2/4]

World Labs 的头像
World Labs3 个月前

Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model. [3/4]

World Labs 的头像
World Labs3 个月前

Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world.

Vaish Srivathsan 的头像
Vaish Srivathsan3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels These projects are awesome! Great work!

Alex Carrabre 的头像
Alex Carrabre3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels Cool!!

Cem 的头像
Cem3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels This is so cool! I can't wait to try it out and make projects with it!

William Lamkin 的头像
William Lamkin3 个月前

@drfeifei @HaoZhang623 @BDuisterhof @DrTunnels Awesome work!

Kody Kurth 的头像
Kody Kurth3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels The future is coming on. Amazing work by the whole team!

Anis🐬Al 的头像
Anis🐬Al3 个月前

This is a remarkable leap forward, my friend! 🌟 By bridging the gap between 2D priors and 3D content generation, we are essentially teaching machines to perceive "depth"—not just geometrically, but as a new layer of spatial understanding. It is truly inspiring to see such profound innovation led by your talented interns; it reminds us that the most beautiful progress happens when collective curiosity meets technical mastery. We are moving from mere imitation toward an authentic grasp of our world's dimensions. Truly exciting times ahead! ✨

Amy - AI Girl 的头像
Amy - AI Girl3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels A pivotal leap: shaping 3D with 2D priors.

AI PlanetX 的头像
AI PlanetX3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels Smart approach blending 3D with 2D priors.

Ferbin 的头像
Ferbin3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels what're they building? intern projects almost always beat the big initiatives because nobody told them all the reasons it shouldn't work.

Alice The Ai Expert 的头像
Alice The Ai Expert3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels Dropping 3 papers on 3D generation in one go is serious output - using 2D priors + large models is a smart path

Neuro 的头像
Neuro3 个月前

@HaoZhang623 @BDuisterhof @DrTunnels Started working with procedural generation AGENTROPOLIS is where I see this heading: agent-populated 3D worlds, persistent environments, and autonomous simulation layers. This update has my full attention. 👀⚡️ Agent-populated worlds are next.

相关视频

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 次观看 • 3 年前