Loading video...
Video Failed to Load
๐ข A Recipe for Generating 3D Worlds From a Single Image ๐ข Our recipe explains how existing generative models can be adapted with minimal training effort to generate 3D worlds from a single input image.
13,970 views โข 1 year ago โขvia X (Twitter)
11 Comments

Our process involves two steps: generating coherent panoramas using a pre-trained inpainting diffusion model and lifting these into 3D with a metric depth estimator.

We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning.

The scene is parameterized by Gaussian Splats and can be explored on a VR headset within a cube with 2m side length.

Our recipe natively extends to text inputs. Here, the prompt is used to first generate the input image.

Check out our project page for more results:

Cinematic pedigree of the highest order meets innovative AAA gameplay in OP3N. Dive into the world of Ready Player One โ Wishlist Now!

@DRozumnyi Looks amazing! Congrats Katja!!

@DRozumnyi Amazing work! Great progress in the world of LWMs and I hope to reimplement this

@DRozumnyi ๐๐

@DRozumnyi Hi Katja, thanks for sharing your work. Is there any plan for open-release of your project?

@DRozumnyi Hey :) We won't release code but the paper should contain all the necessary information to reimplement it
