Загрузка видео...

Не удалось загрузить видео

На главную

📢 A Recipe for Generating 3D Worlds From a Single Image 📢 Our recipe explains how existing generative models can be adapted with minimal training effort to generate 3D worlds from a single input image.

13,970 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

Our process involves two steps: generating coherent panoramas using a pre-trained inpainting diffusion model and lifting these into 3D with a metric depth estimator.

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning.

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

The scene is parameterized by Gaussian Splats and can be explored on a VR headset within a cube with 2m side length.

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

Our recipe natively extends to text inputs. Here, the prompt is used to first generate the input image.

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

Check out our project page for more results:

Фото профиля OPEN
OPEN1 год назад

Cinematic pedigree of the highest order meets innovative AAA gameplay in OP3N. Dive into the world of Ready Player One — Wishlist Now!

Фото профиля Samarth Sinha
Samarth Sinha1 год назад

@DRozumnyi Looks amazing! Congrats Katja!!

Фото профиля Jai Amin
Jai Amin1 год назад

@DRozumnyi Amazing work! Great progress in the world of LWMs and I hope to reimplement this

Фото профиля Philipp Tsipman
Philipp Tsipman1 год назад

@DRozumnyi 👏👏

Фото профиля Chi
Chi1 год назад

@DRozumnyi Hi Katja, thanks for sharing your work. Is there any plan for open-release of your project?

Фото профиля Katja Schwarz
Katja Schwarz1 год назад

@DRozumnyi Hey :) We won't release code but the paper should contain all the necessary information to reimplement it

Похожие видео

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 просмотров • 1 год назад