ะ—ะฐะณั€ัƒะทะบะฐ ะฒะธะดะตะพ...

ะะต ัƒะดะฐะปะพััŒ ะทะฐะณั€ัƒะทะธั‚ัŒ ะฒะธะดะตะพ

ะะฐ ะณะปะฐะฒะฝัƒัŽ

๐Ÿ“ข๐‹๐Ÿ‘๐ƒ๐†: ๐‹๐š๐ญ๐ž๐ง๐ญ ๐Ÿ‘๐ƒ ๐†๐š๐ฎ๐ฌ๐ฌ๐ข๐š๐ง ๐ƒ๐ข๐Ÿ๐Ÿ๐ฎ๐ฌ๐ข๐จ๐ง๐Ÿ“ข #SIGGRAPHAsia We propose a generative diffusion model for 3D Gaussians. Key is a learnt latent space which substantially reduces the complexity of the diffusion process, thus facilitating room-scale scene generation! Great work by Barbara Roessle in with Norman Mรผller, Angela Dai, Lorenzo Porzi, Samuel...

MattNiessner's profile picture

Matthias Niessner

49,592 subscribers

39,529 ะฟั€ะพัะผะพั‚ั€ะพะฒ โ€ข 1 ะณะพะด ะฝะฐะทะฐะด โ€ขvia X (Twitter)

ะšะพะผะผะตะฝั‚ะฐั€ะธะธ: 5

ะคะพั‚ะพ ะฟั€ะพั„ะธะปั Abdullah Hamdi
Abdullah Hamdi1 ะณะพะด ะฝะฐะทะฐะด

Awesome work!

ะคะพั‚ะพ ะฟั€ะพั„ะธะปั Redcrown
Redcrown1 ะณะพะด ะฝะฐะทะฐะด

Cool

ะคะพั‚ะพ ะฟั€ะพั„ะธะปั Station ๐Ÿค–
Station ๐Ÿค–1 ะณะพะด ะฝะฐะทะฐะด

Very cool

ะคะพั‚ะพ ะฟั€ะพั„ะธะปั Lucas Armand
Lucas Armand1 ะณะพะด ะฝะฐะทะฐะด

awesome!

ะคะพั‚ะพ ะฟั€ะพั„ะธะปั Anders Eklund
Anders Eklund1 ะณะพะด ะฝะฐะทะฐะด

Can it be used for 3D medical volumes?

ะŸะพั…ะพะถะธะต ะฒะธะดะตะพ

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior paper page: present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistency issue that existing works encounter. To sculpt geometries that render coherently, we perform score distillation sampling via a view-dependent diffusion model. This 3D prior, alongside several training strategies, prioritizes the geometry consistency but compromises the texture fidelity. We further propose Bootstrapped Score Distillation to specifically boost the texture. We train a personalized diffusion model, Dreambooth, on the augmented renderings of the scene, imbuing it with 3D knowledge of the scene being optimized. The score distillation from this 3D-aware diffusion prior provides view-consistent guidance for the scene. Notably, through an alternating optimization of the diffusion prior and 3D scene representation, we achieve mutually reinforcing improvements: the optimized 3D scene aids in training the scene-specific diffusion model, which offers increasingly view-consistent guidance for 3D optimization. The optimization is thus bootstrapped and leads to substantial texture boosting. With tailored 3D priors throughout the hierarchical generation, DreamCraft3D generates coherent 3D objects with photorealistic renderings, advancing the state-of-the-art in 3D content generation.

AK

161,530 ะฟั€ะพัะผะพั‚ั€ะพะฒ โ€ข 2 ะปะตั‚ ะฝะฐะทะฐะด

LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models paper page: github: Recent advancements in text-to-image generation with diffusion models have yielded remarkable results synthesizing highly realistic and diverse images. However, these models still encounter difficulties when generating images from prompts that demand spatial or common sense reasoning. We propose to equip diffusion models with enhanced reasoning capabilities by using off-the-shelf pretrained large language models (LLMs) in a novel two-stage generation process. First, we adapt an LLM to be a text-guided layout generator through in-context learning. When provided with an image prompt, an LLM outputs a scene layout in the form of bounding boxes along with corresponding individual descriptions. Second, we steer a diffusion model with a novel controller to generate images conditioned on the layout. Both stages utilize frozen pretrained models without any LLM or diffusion model parameter optimization. We validate the superiority of our design by demonstrating its ability to outperform the base diffusion model in accurately generating images according to prompts that necessitate both language and spatial reasoning. Additionally, our method naturally allows dialog-based scene specification and is able to handle prompts in a language that is not well-supported by the underlying diffusion model.

AK

83,681 ะฟั€ะพัะผะพั‚ั€ะพะฒ โ€ข 3 ะปะตั‚ ะฝะฐะทะฐะด

Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution paper page: Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and quality. Extensive experiments show that Upscale-A-Video surpasses existing methods in both synthetic and real-world benchmarks, as well as in AI-generated videos, showcasing impressive visual realism and temporal consistency.

AK

32,849 ะฟั€ะพัะผะพั‚ั€ะพะฒ โ€ข 2 ะปะตั‚ ะฝะฐะทะฐะด