Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors. These projects were led by our incredible interns Hao Zhang @BDuisterhof @DrTunnels [1/4]

122,868 Aufrufe • vor 3 Monaten •via X (Twitter)

15 Kommentare

Profilbild von Fei-Fei Li
Fei-Fei Livor 3 Monaten

@BenMildenhall @HaoZhang623 @BDuisterhof @DrTunnels Very fun to work with our awesome WL interns!😍🌐

Profilbild von World Labs
World Labsvor 3 Monaten

World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image. [2/4]

Profilbild von World Labs
World Labsvor 3 Monaten

Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model. [3/4]

Profilbild von World Labs
World Labsvor 3 Monaten

Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world.

Profilbild von Vaish Srivathsan
Vaish Srivathsanvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels These projects are awesome! Great work!

Profilbild von Alex Carrabre
Alex Carrabrevor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels Cool!!

Profilbild von Cem
Cemvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels This is so cool! I can't wait to try it out and make projects with it!

Profilbild von William Lamkin
William Lamkinvor 3 Monaten

@drfeifei @HaoZhang623 @BDuisterhof @DrTunnels Awesome work!

Profilbild von Kody Kurth
Kody Kurthvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels The future is coming on. Amazing work by the whole team!

Profilbild von Anis🐬Al
Anis🐬Alvor 3 Monaten

This is a remarkable leap forward, my friend! 🌟 By bridging the gap between 2D priors and 3D content generation, we are essentially teaching machines to perceive "depth"—not just geometrically, but as a new layer of spatial understanding. It is truly inspiring to see such profound innovation led by your talented interns; it reminds us that the most beautiful progress happens when collective curiosity meets technical mastery. We are moving from mere imitation toward an authentic grasp of our world's dimensions. Truly exciting times ahead! ✨

Profilbild von Amy - AI Girl
Amy - AI Girlvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels A pivotal leap: shaping 3D with 2D priors.

Profilbild von AI PlanetX
AI PlanetXvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels Smart approach blending 3D with 2D priors.

Profilbild von Ferbin
Ferbinvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels what're they building? intern projects almost always beat the big initiatives because nobody told them all the reasons it shouldn't work.

Profilbild von Alice The Ai Expert
Alice The Ai Expertvor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels Dropping 3 papers on 3D generation in one go is serious output - using 2D priors + large models is a smart path

Profilbild von Neuro
Neurovor 3 Monaten

@HaoZhang623 @BDuisterhof @DrTunnels Started working with procedural generation AGENTROPOLIS is where I see this heading: agent-populated 3D worlds, persistent environments, and autonomous simulation layers. This update has my full attention. 👀⚡️ Agent-populated worlds are next.

Ähnliche Videos

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 Aufrufe • vor 3 Jahren