Introducing ๐ฆ๐๐ฟ๐๐ถ๐๐ฎ๐๐ฒ๐ป๐๐ง (SIGGRAPH Asia 2025) โ a high-quality 3D... diffusion model that explicitly models object articulation, paving the way for richer, more realistic assets in embodied AI and simulation: โ Generates fully articulated 3D objects โ Physically plausible joints & motion โ High-fidelity 3D Gaussian appearance โ Supports generation from a single real image arXiv: Project: Code (coming soon):show more

Xingang Pan
11,520 ๆฌก่ง็ โข 9 ไธชๆๅ
Static 3D generation isn't enough. We need assets ready... for animation. Our new #SIGGRAPH work, AniGen, takes a single image and generates the 3D shape, skeleton, and skinning weights all at once. Code is fully open-sourced! Kudos to Yihua and VAST AI Research ๐งต(1/4)show more

Yanpei Cao
145,188 ๆฌก่ง็ โข 4 ไธชๆๅ
๐ Introducing GenLit โ Reformulating Single-Image Relighting as Video... Generation! We leverage video diffusion models to perform realistic near-field relighting from just a single imageโNo explicit 3D reconstruction or ray tracing required! No intermediate graphics buffers, directly in the pixel space! ๐ Dive into the paper: ๐ฅ Project page & demos: ๐ Code coming soon! #GenerativeAI #ComputerVision #Relighting #DiffusionModels #Graphics ๐งต 1/5show more

Haven Feng
22,474 ๆฌก่ง็ โข 1 ๅนดๅ
As announced in partnership with NVIDIA at CES, weโre... excited to introduce Stable Point Aware 3D (SPAR3D), setting a new standard in 3D generation. Ideal for running on NVIDIA RTX AI PCs, SPAR3D enables real-time editing and complete structure generation of 3D objects from a single image in under a second. You can download the weights on Hugging Face and code on GitHub, or access the model through the Stability AI API. Learn more here: (1/3)show more

Stability AI
181,554 ๆฌก่ง็ โข 1 ๅนดๅ
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): โข We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. โข We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. โข To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. โข Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,062 ๆฌก่ง็ โข 1 ๅนดๅ
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... โข First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. โข Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. โข Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasksโand bridging image space and 3D spaceโthrough the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256ร spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,849 ๆฌก่ง็ โข 1 ๅนดๅ
Create a 3D model from a single image, set... of images or a text prompt in < 1 minute ๐ฎโ๐จ This new AI paper called CAT3D shows us that itโll keep getting easier to produce 3D models from 2D images โ whether itโs a sparser real world 3D scan (a few photos instead of hundreds) or your favorite 2D image generator like Midjourney (just an image). How does this magic work? โThis architecture is similar to video diffusion models, but with camera pose embeddings for each image instead of time embeddings. The generated views are passed into a robust 3D reconstruction pipeline to create the 3D representation (Zip-NeRF or 3DGS)โshow more

Bilawal Sidhu
92,867 ๆฌก่ง็ โข 2 ๅนดๅ
With Hunyuan3D World Model 1.0 now released and open-sourced,... we're excited to showcase the technical highlights behind this impressive innovation: โ 360ยฐ Panoramic Generation: Creates complete, immersive โworld scenesโ, far beyond localized views. โ Explorable 3D Scene Generation: Generates diverse, spatially consistent 3D worlds from text/image for truly immersive exploration. โ Interactive/Editable: Achieves separation of foreground objects, background terrain, ground, and sky, for seamless secondary editing. โ Exportable Mesh: Generated scenes can be exported as 3D meshes for direct import into mainstream game engines and modeling software. โ Industry-Leading SOTA Evaluation: Surpasses state-of-the-art open-source models in generation quality. As the industry's first open-source model for physical simulation and explorable world generation, Hunyuan3D World Model 1.0 aims to foster a collaborative community ecosystem with developers and enthusiasts. โจ Try it now: ๐ค Hugging Face:show more

Tencent Hy
23,203 ๆฌก่ง็ โข 1 ๅนดๅ
You can't 3D reconstruct glass from images... ...WRONG! Thanks... for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AIshow more

Jonathan Stephens
17,712 ๆฌก่ง็ โข 8 ไธชๆๅ
There are way too many AI 3D models claiming... to be the best. So weโre testing them head-to-head. Tripo P1 vs Meshy V6 vs Hunyuan V3.1 Pro vs Rodin 2.5 Same image. Same prompt. No cherry-picking, only 1 generation. This weekโs test: a shiny metal helmet. reflections, materials, geometry, and texture accuracy. Next up: characters, objects, environments, and more. By the end, we want one answer: Which AI 3D model should you actually use for your next game, website, or project?show more

mint
20,027 ๆฌก่ง็ โข 20 ๅคฉๅ
2023 was the year of AI avatars 2024 was... the year of AI photos 2025 was the year of AI videos And I think it's becoming clear now that 2026 will be the year of AI world models Fully interactive explorable 3d worlds generated from one or multiple 2d images or a prompt In turn these 2d images can then be generated by AI too So soon you can generate fully explorable virtual 3d worlds based on your own imagination Next will be figuring out how to make those worlds interactive This is World Labs (unaffiliated, but I like it) As always a lot of big AI model companies are now working on the same thing: 3d world models, only World Labs has a real properly working demo (for now) Very exciting time again!show more

@levelsio
584,511 ๆฌก่ง็ โข 11 ไธชๆๅ
Video diffusion models have strong implicit representations of 3D... shape, material, and lighting, but controlling them with language is cumbersome, and control is critical for artists and animators. GenLit connects these implicit representations with a continuous 5D control signal describing the direction and intensity of a point light source. This enables single-image near-field relighting of an image using a video diffusion model. We use a ControlNet-like approach and show that, with a small amount of synthetic data, GenLit generalizes to complex real-world images. Given a single image and the 5D lighting signal, GenLit creates a video of a moving light source that is inside the scene. It moves around and behind scene objects, producing effects such as shading, cast shadows, secularities, and interreflections with a realism that is hard to obtain with traditional inverse rendering methods. GenLit shows that it is possible to get continuous control over implicit physical processes within a video model. I think this is just the beginning and promises to make such models much more practical for creators. Shrisha Bharadwaj will present today at SIGGRAPH Asia Room: S423/S424, Level 4 @ 13:50 on 15 of Dec.show more

Michael Black
22,182 ๆฌก่ง็ โข 8 ไธชๆๅ
So, today we have fast SDF sculpting + real-time... AI in Unbound Loop. Old news๐ฅฑ Coming up next: -quad-view generation from sculpted geometry -image tweaks via nano๐ -tripo HD & Low-Poly 3D generation There's more in the upcoming release, but these three deserve a closer look: Quad-View Generation Most platforms offer some version of this, but Loop has a key advantage, your sculpted model is the reference. That means less guessing from the generator. Though itโs not 100% foolproof, like any AI I guess? (I should stop stating the obvious every time). Image Tweaks via Chat Select any generated image and ask for fixes or changes in real time. Works great on unintended quad-view hallucinations, but also handy for quick iterations. Swapping colors, tweaking details, removing elements. Tripo 3D generator Especially the low-poly model, it consistently delivered fantastic game-ready topology when we tried it with our real-time AI output. And it's super fast.show more

Andrea Intg.
15,132 ๆฌก่ง็ โข 1 ไธชๆๅ
WeatherEdit: Controllable Weather Editing with 4D Gaussian Field Contributions:... 1. Based on our analysis of weather editing characteristics, we introduce WeatherEdit, a comprehensive and efficient framework for realistic and controllable weather generation. Compared with existing methods that focus on either background editing or static weather effects, a progressive 2D-to-4D transformation process in WeatherEdit enhances adaptability across a wider range of scenarios. 2. We introduce an all-in-one adapter to enable a diffusion model for multi-weather (snowy, rainy, and fog) synthesis, along with a Temporal-View attention to ensure consistent editing across multi-frame and multi-view. 3. We design a 4D Gaussian field for weather particle modeling, enabling plausible simulation of raindrops, snowflakes, and fog with controllable severity. 4. We demonstrate WeatherEditโs effectiveness in generating realistic, consistent, and controllable weather effects in 3D driving scenes, showcasing its applicability to real-world scenarios.show more

MrNeRF
10,691 ๆฌก่ง็ โข 1 ๅนดๅ
3D Gaussian Splatting for Real-Time Radiance Field Rendering paper... page: Radiance Field methods have recently revolutionized novel-view synthesis of scenes captured with multiple photos or videos. However, achieving high visual quality still requires neural networks that are costly to train and render, while recent faster methods inevitably trade off speed for quality. For unbounded and complete scenes (rather than isolated objects) and 1080p resolution rendering, no current method can achieve real-time display rates. We introduce three key elements that allow us to achieve state-of-the-art visual quality while maintaining competitive training times and importantly allow high-quality real-time (>= 30 fps) novel-view synthesis at 1080p resolution. First, starting from sparse points produced during camera calibration, we represent the scene with 3D Gaussians that preserve desirable properties of continuous volumetric radiance fields for scene optimization while avoiding unnecessary computation in empty space; Second, we perform interleaved optimization/density control of the 3D Gaussians, notably optimizing anisotropic covariance to achieve an accurate representation of the scene; Third, we develop a fast visibility-aware rendering algorithm that supports anisotropic splatting and both accelerates training and allows realtime rendering. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets.show more

AK
633,674 ๆฌก่ง็ โข 3 ๅนดๅ
Chop the gradients โ๏ธ! We found that truncating decoder... gradients in latent video diffusion to a fixed window allows us to finetune on videos with pixel-wise perceptual losses without running out of memory. Pixel losses have been essential for image generation and reconstruction, but until now, they haven't scaled to long-duration, high-resolution video diffusion due to recursive activation accumulation in causal decoders, leading to OOM during training ๐ฅ๐. Project: Video diffusion models can do a lot more ๐ when you can backprop the decoder! Post-process neural rendered scenes, super-resolve videos, harmonize lighting in controlled synthetic driving scenes, and inpaint videos โ all in a single step โก with a quick finetune from a standard diffusion model.show more

Felix Heide
28,399 ๆฌก่ง็ โข 4 ไธชๆๅ
Introducing Kaleido๐ฎ from AI at Meta โ a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design thatโs simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of โrealities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. ๐ Explore more results and paper:show more

Shikun Liu
22,442 ๆฌก่ง็ โข 11 ไธชๆๅ