Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

📢 A Recipe for Generating 3D Worlds From a Single Image 📢 Our recipe explains how existing generative models can be adapted with minimal training effort to generate 3D worlds from a single input image.

13,970 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

Our process involves two steps: generating coherent panoramas using a pre-trained inpainting diffusion model and lifting these into 3D with a metric depth estimator.

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning.

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

The scene is parameterized by Gaussian Splats and can be explored on a VR headset within a cube with 2m side length.

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

Our recipe natively extends to text inputs. Here, the prompt is used to first generate the input image.

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

Check out our project page for more results:

OPEN profil fotoğrafı
OPEN1 yıl önce

Cinematic pedigree of the highest order meets innovative AAA gameplay in OP3N. Dive into the world of Ready Player One — Wishlist Now!

Samarth Sinha profil fotoğrafı
Samarth Sinha1 yıl önce

@DRozumnyi Looks amazing! Congrats Katja!!

Jai Amin profil fotoğrafı
Jai Amin1 yıl önce

@DRozumnyi Amazing work! Great progress in the world of LWMs and I hope to reimplement this

Philipp Tsipman profil fotoğrafı
Philipp Tsipman1 yıl önce

@DRozumnyi 👏👏

Chi profil fotoğrafı
Chi1 yıl önce

@DRozumnyi Hi Katja, thanks for sharing your work. Is there any plan for open-release of your project?

Katja Schwarz profil fotoğrafı
Katja Schwarz1 yıl önce

@DRozumnyi Hey :) We won't release code but the paper should contain all the necessary information to reimplement it

Benzer Videolar

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 görüntüleme • 1 yıl önce