Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Excited to share our work on Neural Assets: a new method for enabling 3D asset-level control in image diffusion models – scalable & without any 3D inductive biases. Neural Assets goes beyond text or pixel-based control & provides an interface inspired by 3D graphics tools. 🧵

97,924 görüntüleme • 2 yıl önce •via X (Twitter)

9 Yorum

Thomas Kipf profil fotoğrafı
Thomas Kipf2 yıl önce

Paper: Website: Neural Assets enables a range of 3D editing capabilities for individual or multiple assets: translation, rotation, rescaling, transfer across scenes, and control over the scene background.

Thomas Kipf profil fotoğrafı
Thomas Kipf2 yıl önce

Assets extracted from one scene and placed into a different scene or background naturally adapt to lighting conditions and other environmental factors. At night or in rainy conditions, cars even turn their lights/headlights on!

Thomas Kipf profil fotoğrafı
Thomas Kipf2 yıl önce

Neural Assets are extracted from raw video frames with the help of 2D or 3D boxes. The key to make it work is to extract appearance and pose representations from *different* frames, which results in disentanglement and thus controllability. The entire model is trained/fine-tuned end-to-end jointly with a pre-trained image generation model (here: Stable Diffusion 2.1) simply by replacing the text token sequence in the base text-to-image model with a Neural Assets token sequence. At test time, we can compose scenes by combining multiple Neural Assets and a neural representation of the scene background, while ensuring appearance consistency (to a large extent) both for individual assets as well as the overall scene.

Thomas Kipf profil fotoğrafı
Thomas Kipf2 yıl önce

Check out our paper ( and our website ( for a lot more details, results, and current limitations / failure modes. Neural Assets is the result of @Dazitu_616's outstanding work as a student researcher in our team, working with a set of fantastic collaborators: @YuliaRubanova, @RishabhKabra, @drewAhudson, @yusufaytar, @vansteenkiste_s, @KelseyRAllen; advised by @igilitschenski. Starting with Slot Attention in 2020, we have pursued this research direction over the past four years (SAVi, OSRT, DORSal & many other works). I couldn't be more excited about this latest result and the potential for this class of methods to enable new creative control capabilities for image generation models and beyond.

Yulia Rubanova profil fotoğrafı
Yulia Rubanova2 yıl önce

Super excited to be part of this work. Using the same interface of Neural Assets, we can get a rich set of controls over the objects and seamlessly blend the objects into the environment, with appropriate lighting and shadows. Amazing work, @Dazitu_616!

Ziyi Wu profil fotoğrafı
Ziyi Wu2 yıl önce

Thank you, Thomas! It's been an awesome experience working with you and all the Google folks. I will definitely miss this Student Researcher journey!

Omri Kaduri profil fotoğrafı
Omri Kaduri2 yıl önce

That's really great to see the progress you are making on object-centric representations. Does this model and code will be released?

Nate Codes profil fotoğrafı
Nate Codes1 yıl önce

When people can do this inside of the physical neural asset space things I'm going to be lost in AR/VR! I love your work and it has tons of implications for my work.

Thomas Kipf profil fotoğrafı
Thomas Kipf1 yıl önce

Thanks, Nate! Lots of work still to be done by the machine learning community before an approach like this becomes widely usable, but I’m personally really excited about this future.

Benzer Videolar

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 görüntüleme • 3 yıl önce