Loading video...

Video Failed to Load

Go Home

Can we use video diffusion to generate 3D scenes? ๐–๐จ๐ซ๐ฅ๐๐„๐ฑ๐ฉ๐ฅ๐จ๐ซ๐ž๐ซ (#SIGGRAPHAsia25) creates fully-navigable scenes via autoregressive video generation. Text input -> 3DGS scene output & interactive rendering! ๐ŸŒ ๐Ÿ“ฝ๏ธ

30,883 views โ€ข 10 months ago โ€ขvia X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

๐Ÿšจ SIGGRAPH Asia 2025 Paper Alert ๐Ÿšจ โžก๏ธPaper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes ๐ŸŒŸFew pointers from the paper ๐ŸŽฏGenerating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. ๐ŸŽฏ To this end, authors of this paper proposed โ€œWorldExplorerโ€, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. ๐ŸŽฏThey initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. ๐ŸŽฏThen, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. ๐ŸŽฏConcretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. ๐ŸŽฏTheir novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. ๐ŸŽฏFinally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. ๐ŸŽฏCompared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. ๐ŸŽฏThey believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. ๐ŸขOrganization: TU Mรผnchen ๐Ÿง™Paper Authors: Manuel-Andreas Schneider, Lukas Hรถllein , Matthias Niessner ๐Ÿ“ Read the Full Paper here: ๐Ÿ—‚๏ธ Project Page: ๐Ÿง‘โ€๐Ÿ’ป Code: ๐ŸŽฅ Be sure to watch the attached Technical Summary Video - Sound on ๐Ÿ”Š๐Ÿ”Š Find this Valuable ๐Ÿ’Ž ? โ™ป๏ธQT and teach your network something new Follow me ๐Ÿ‘ฃ, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 views โ€ข 10 months ago

๐Ÿ“ข๐Ÿ“ข ๐๐ž๐ซ๐œ๐‡๐ž๐š๐: ๐๐ž๐ซ๐œ๐ž๐ฉ๐ญ๐ฎ๐š๐ฅ ๐‡๐ž๐š๐ ๐Œ๐จ๐๐ž๐ฅ ๐Ÿ๐จ๐ซ ๐’๐ข๐ง๐ ๐ฅ๐ž-๐ˆ๐ฆ๐š๐ ๐ž ๐Ÿ‘๐ƒ ๐‡๐ž๐š๐ ๐‘๐ž๐œ๐จ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง & ๐„๐๐ข๐ญ๐ข๐ง๐ ๐Ÿ“ข๐Ÿ“ข PercHead reconstructs realistic 3D heads from a single image and enables disentangled 3D editing via geometric controls and style inputs from images or text. At its core is a generalized 3D head decoder trained with perceptual supervision from DINOv2 and SAM 2.1. We find that our new perceptual loss formulation improves reconstruction fidelity compared to commonly-used methods such as LPIPS. Our trained reconstruction model is able to generate 3D-consistent heads from a single input image. Even with challenging side-view inputs, the model robustly infers missing regions for a coherent, high-fidelity output. In addition, our architecture seamlessly adapts to downstream tasks: by swapping the encoder, we can transform the model into a disentangled 3D editing pipeline. In this scenario, we can control geometry through - potentially hand-drawn - segmentation maps, and condition style via image or text prompt. We also provide an interactive GUI to enable the exploration of our editing pipeline. ๐ŸŒ ๐Ÿ“ฝ๏ธ Great work by Antonio Oroz and Tobias Kirschstein

Matthias Niessner

18,855 views โ€ข 9 months ago