Loading video...

Video Failed to Load

Go Home

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them...

34,440 views • 1 year ago •via X (Twitter)

10 Comments

Bilawal Sidhu's profile picture
Bilawal Sidhu1 year ago

Check out the paper / project here -- pretty nice thread on the methodology:

Mobile Scanner's profile picture
Mobile Scanner1 year ago

Scan any documents, convert images into text, PDF files, etc. 👍

Michael Gold's profile picture
Michael Gold1 year ago

Very interesting approach: 2D -> 2.5D -> 3D -> 2D render

Captain HaHaa's profile picture
Captain HaHaa1 year ago

That's wild mate! What a concept 🤯

Creative Cheat Code 🕹️'s profile picture
Creative Cheat Code 🕹️1 year ago

I've always thought difusion models are just missing a layer of intelligence about physics. I'm sure we'll get to a point of combining these two worlds

Eriiiiiiiick's profile picture
Eriiiiiiiick1 year ago

Big B, where do you find this stuff?

Bilawal Sidhu's profile picture
Bilawal Sidhu1 year ago

often on the TL, but many times researchers are kind enough new papers with me. keep em coming!

Supreme's profile picture
Supreme1 year ago

Woah

Lachlan Phillips exo/acc 👾's profile picture
Lachlan Phillips exo/acc 👾1 year ago

Temporally consistent SDXL is the real game changer

Clyde DeSouza's profile picture
Clyde DeSouza1 year ago

This is what Canoma could have evolved to...if Adobe hadn't killed it

Related Videos

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior paper page: present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistency issue that existing works encounter. To sculpt geometries that render coherently, we perform score distillation sampling via a view-dependent diffusion model. This 3D prior, alongside several training strategies, prioritizes the geometry consistency but compromises the texture fidelity. We further propose Bootstrapped Score Distillation to specifically boost the texture. We train a personalized diffusion model, Dreambooth, on the augmented renderings of the scene, imbuing it with 3D knowledge of the scene being optimized. The score distillation from this 3D-aware diffusion prior provides view-consistent guidance for the scene. Notably, through an alternating optimization of the diffusion prior and 3D scene representation, we achieve mutually reinforcing improvements: the optimized 3D scene aids in training the scene-specific diffusion model, which offers increasingly view-consistent guidance for 3D optimization. The optimization is thus bootstrapped and leads to substantial texture boosting. With tailored 3D priors throughout the hierarchical generation, DreamCraft3D generates coherent 3D objects with photorealistic renderings, advancing the state-of-the-art in 3D content generation.

AK

161,530 views • 2 years ago

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 views • 3 years ago

So these researchers figured out you can basically hallucinate 3D cities into existence using just satellite photos & a diffusion model. The problem's pretty straightforward: satellites only see rooftops. Building facades? Invisible. Street-level detail? Doesn't exist. But people want flyable 3D environments, which means you need all that occluded geometry. When I worked on google maps photogrammetry, we could only use satellite-based 3D for isolated stuff like the pyramids - anything city-scale required airplane flyovers. Which is fine until you hit aerial-denied regions where you literally can't fly. Huge chunks of the world just unavailable. Their trick is honestly kind of beautiful. They train gaussian splats on satellite views, but as it descends toward ground level, the renders turn to absolute garbage - artifacts everywhere. Instead of fighting this, they just treat those nightmare renders as the input to a diffusion model. Basically - "hey FLUX, fix this mess." Then here's where it gets clever: they generate multiple diffusion samples per view instead of committing to one. Because any single denoising path is probably wrong in 3D space, but if you generate a couple and let the GS optimization find consensus across them, you get actual geometric consistency. They do this in episodes, curriculum style - start high, gradually descend (hence the name Skyfall-GS!). With each iteration the ground-level views get less fucked. By the end you've got real-time flyable cities that look surprisingly real, and the geometry still matches the satellite input. No 3D training data. No street-level photos. Just satellites + diffusion doing what it does best - filling in the blanks. It's like neural scene completion but actually practical, and it unlocks basically the entire world.

Bilawal Sidhu

241,899 views • 9 months ago

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 views • 8 months ago

Here's a devlog made by an anonymous Chinese fan replicating the surprisingly brand new technique that I developed for detecting asteroids which wound up being so powerful that it can easily track Stealth Fighters from over 100km away even when it’s only using three $30 webcams as sensors meaning it easily outperforms all modern stealth tracking techniques in precision, range and cost. And while this demo is using optical light, this same technique which I call pixel motion to voxel projection, can be used interchangeably with thermal infrared cameras to work at night and also majorly boosts the effectiveness of radar allowing you to track fighters much more effectively through clouds and over the horizon. This technique will also always eventually give the exact location of the target even if the image is blurry as those blurs will always average out from the different perspectives into revealing the precise location of the target in the voxel grid. There is definitely a Mandela effect with this technique as it feels as though it should already exist, especially because at first as it sounds like it is performing triangulation (which has existed for years and is what we do for mocap and tennis ball tracking). But triangulation is entirely separate to this as triangulations only works if you have already identified where the ball is in a 2D image because you’re able to rely on being able to use at least 2 separate high quality cameras which are much closer to the ball making the ball’s apparent size much much bigger and therefore gives you hundreds of pixels to work with which makes it much easier to use object recognition techniques to recognize where it is in the image aka in 2D and then you’re just using the other cameras view to project out lines which intersect in 3D to find out where the ball is in 3D. The major difference is that pixel motion to voxel projection allows you to find where the object is in 3D without having already found it in 2D which is an unbelievable difference as it allows you to use much lower quality cameras together to accumulate data together into 3D space. If this seem like it doesn’t mean much then what it actually means is that you don’t understand what I’m saying as what I’m saying means a LOT in practical terms as it means you go from having to use an imaging system that has to be able to image the object to the point that it is over a hundred total pixels in surface area to have enough data to recognize it to instead be able to use something that is only images the object to be 1 pixel in surface area and only changes the brightness value by 1 value every now and then. I’d recommend an amazing video by DST studios called “Lowlight cameras can’t defeat stealth” if you want a great video which goes over the difficulty of even using telescopes to recognize stealth fighters and why this is so impressive compared to other techniques and ironically it is what inspired me to realize the asteroid tracker I was working on actually could do this. Which brings me to the point that if this wasn’t a new technique then not only would there be at least one example of an asteroid survey that points distant telescopes at the same place at the same time in order to be able to add the light together to detect asteroids which as I was shocked to learn isn’t a thing despite the fact that it would make detecting asteroids trivial by comparison to modern 2D imaging while also having no impact on the normal scientific operations of those surveys other than small changes to scheduling. But there would also be an example of a drone tracker that uses this instead of using the aforementioned high quality zoomable telescope which has to be able to zoom in close enough to be able to recognize a drone. If you want to tell me that this is something that already exists give me an exact example of a product that uses it, not the general outline of a concept that you think it is, the actual product and then also tell me the asteroid survey that uses distant telescopes that point at the exact same place at the exact same time because I can guarantee that if you google what you think uses this you won’t even find the steps of subtracting the images from each other to get motion and will definitely not get the added step of projecting that motion into a voxel grid (It would blow your mind if you found out how Xbox kinect cameras work.) Also I want to make it clear, I’m not saying you should just use web cams to do this, I’m just using them as an example to show you the power of this in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar. Pretty much all of the problems you could think of for this are incredibly easy to overcome if you apply even a small amount of brainpower into fixing the problem. And yes, this gives you the exact location down to the meter of whatever you are tracking even if the image is blurry as those blurs will always average out to the exact location down to the meter in the voxel grid. Which is what makes this technique so powerful since the cost of adding each camera to The network grows linearly while the rate at which each camera gives more information grows exponentially due to the increasing unlikeliness of all of them having more movement in the same place. And given the size of the cameras it really wouldn’t be that hard to hide and network these cameras together in other countries and on sea buoys to know where planes are everywhere in the world. Which brings me to the point that I personally really don’t care about the military uses of this technology, if all it could do is precisely track stealth fighters then I wouldn’t have cared enough to work on it, I could have used any of the many other life saving techniques as the subject of the video, stealth fighters just sounds the most clickable and the scale of the problem is more intuitive to most people and if I did use any of those as subjects for the demo it would inevitably result in the stealth fighter technique being figured out anyway and all of the other uses are so useful that I don't think anyone would reasonably complain about the upside. The real purpose of this video is that since this is a new technique that hasn’t been used to detect stealth fighters despite the billions we have spent on that, then what else can you apply this to that could go on to improve billions of people’s lives that you or others are working on. For example this also allows you to majorly improve the effectiveness of cryo electron microscopy and CT scanners. This part also is kind of hard to explain as it also sounds like it exists but again, when you look through all of the places where you think it is being used you will find that it wasn’t. What I’m saying here isn’t that this is a Radon transform or gaussian splat or whatever, I’m saying that this is able to get new information that wasn’t being accessed before due to the added information about depth you get from the correlation of movement between each perspective which adds to the information that you already have. This allows you to directly subtract foreground and background objects as well as noise faster than you would be able to before and works better than super resolution for your images since super resolution won’t remove foreground and background objects like this does and instead just scales up target, foreground and background objects indiscriminately. And while with enough data Radon transforms or other scanning techniques would eventually get you a correct answer this will get you there a lot faster since those are mostly averaging techniques which average out noise whereas this gets you the ability to directly subtract noise. I’m not expecting you to think that this would do anything but if you try it for yourself you will find that it does majorly improve your ability to perform 3d scans. Again, cryo EM is a field where you would expect this technique to exist but when you look through all the papers on the topic there is no mention of tilting the grid slightly in order to be able to change your perspective slightly on the order of the feature size (if you tilt the grid then you only need precision on the order of an arc minute to do this) and doing multiple exposures from multiple different known tilts and then using those difference images to correlate depth from motion. In fact, in cryo EM you would normally want to do the opposite of this and have your exposures all taken from the same grid angle and just use the variations in how many of the same proteins are oriented in order to be able to scan them for a 3D model but this will generate you far more data faster. There is so much information that I can’t really explain in text so if you have any questions such as why this hasn’t been made before then they will most likely be answered in the video I originally posted which I have added to the end of the first Devlog for your convenience. And again, pretty much all of the problems with the technique can be fixed with a little bit of brainpower, in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar.

ConsistentlyInconsistent

50,687 views • 11 months ago

🔴 Finally! NVIDIA has finally made the code for Neuralangelo public! It has the ability to transform any video into a highly detailed 3D environment, and it's a technology related to but DIFFERENT from NeRF. 💡 Here's how it works: It takes a 2D video as input, showing an object, monument, building, landscape, etc., from various perspectives and analyzes details such as depth, size, and the shapes of objects. From this, the AI sketches an initial 3D model, similar to how an artist molds a figure. This representation is then refined to highlight more details, just as an artist would make the final touches when sculpting. The result is a 3D environment/model, perfect for use in any environment. Imagine the applications it will have for video games, cinema, virtual environments, VR, and more! 📽️🎮 💡 More details: A year ago, an article was presented on a groundbreaking technique called NVIDIA's Instant NeRF. This technique turns images into stunning 3D scenes in a short time, ideal for creating realistic models for video games and other applications. Although Instant NeRF had a lot of potential, the generated models were not perfect and often lacked detailed structures, appearing somewhat cartoonish. A year on, NVIDIA releases a new technique based on Instant NeRF, named Neuralangelo. This enhances the fidelity of surface structures. While NeRF reconstructs real objects in virtual environments from images or videos, Instant NeRF speeds up this process, and Neuralangelo further improves the quality, making the generated objects appear even more realistic when examined up close. Neuralangelo improves Instant NeRF's approach in two key ways related to the hash grid encoding technique: 1⃣ Numerical gradients have been used to compute higher-order derivatives as a smoothing operation. This optimizes the "hash grid" encoding using numerical rather than analytical gradients, providing a smoother input to the network that produces the 3D model. 2⃣ A "coarse-to-fine" optimization has been implemented in the hash grids to control different levels of detail. That is, they first focus on a smoothed version of the scene, and then refine it with more detailed updates. Well, as Arthur C. Clarke said, "Any sufficiently advanced technology is indistinguishable from magic."

Javi Lopez ⛩️

689,180 views • 3 years ago

The most interesting part for me is where Andrej Karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforcement learning is terrible. It just so happens that everything that we had before is much worse.” So what do humans do instead? > “The book I’m reading is a set of prompts for me to do synthetic data generation. It's by manipulating that information that you actually gain that knowledge. We have no equivalent of that with LLMs; they don't really do that.” > “I'd love to see during pretraining some kind of a stage where the model thinks through the material and tries to reconcile it with what it already knows. There's no equivalent of any of this. This is all research.” Why can’t we just add this training to LLMs today? > “There are very subtle, hard to understand reasons why it's not trivial. If I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' You could try, but the model will actually get much worse if you continue trying.” > “Say we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same.” > “You're not getting the richness and the diversity and the entropy from these models as you would get from humans. How do you get synthetic data generation to work despite the collapse and while maintaining the entropy? It is a research problem.” How do humans get around model collapse? > “These analogies are surprisingly good. Humans collapse during the course of their lives. Children haven't overfit yet. They will say stuff that will shock you. Because they're not yet collapsed. But we [adults] are collapsed. We end up revisiting the same thoughts, we end up saying more and more of the same stuff, the learning rates go down, the collapse continues to get worse, and then everything deteriorates.” In fact, there’s an interesting paper arguing that dreaming evolved to assist generalization, and resist overfitting to daily learning - look up The Overfitted Brain by Erik Hoel. I asked Karpathy: Isn’t it interesting that humans learn best at a part of their lives (childhood) whose actual details they completely forget, adults still learn really well but have terrible memory about the particulars of the things they read or watch, and LLMs can memorize arbitrary details about text that no human could but are currently pretty bad at generalization? > “[Fallible human memory] is a feature, not a bug, because it forces you to only learn the generalizable components. LLMs are distracted by all the memory that they have of the pre-trained documents. That's why when I talk about the cognitive core, I actually want to remove the memory. I'd love to have them have less memory so that they have to look things up and they only maintain the algorithms for thought, and the idea of an experiment, and all this cognitive glue for acting.”

Dwarkesh Patel

1,051,605 views • 10 months ago