Loading video...

Video Failed to Load

Go Home

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them...

34,440 views • 1 year ago •via X (Twitter)

10 Comments

Bilawal Sidhu's profile picture
Bilawal Sidhu1 year ago

Check out the paper / project here -- pretty nice thread on the methodology:

Mobile Scanner's profile picture
Mobile Scanner1 year ago

Scan any documents, convert images into text, PDF files, etc. 👍

Michael Gold's profile picture
Michael Gold1 year ago

Very interesting approach: 2D -> 2.5D -> 3D -> 2D render

Captain HaHaa's profile picture
Captain HaHaa1 year ago

That's wild mate! What a concept 🤯

Creative Cheat Code 🕹️'s profile picture
Creative Cheat Code 🕹️1 year ago

I've always thought difusion models are just missing a layer of intelligence about physics. I'm sure we'll get to a point of combining these two worlds

Eriiiiiiiick's profile picture
Eriiiiiiiick1 year ago

Big B, where do you find this stuff?

Bilawal Sidhu's profile picture
Bilawal Sidhu1 year ago

often on the TL, but many times researchers are kind enough new papers with me. keep em coming!

Supreme's profile picture
Supreme1 year ago

Woah

Lachlan Phillips exo/acc 👾's profile picture
Lachlan Phillips exo/acc 👾1 year ago

Temporally consistent SDXL is the real game changer

Clyde DeSouza's profile picture
Clyde DeSouza1 year ago

This is what Canoma could have evolved to...if Adobe hadn't killed it

Related Videos

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior paper page: present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistency issue that existing works encounter. To sculpt geometries that render coherently, we perform score distillation sampling via a view-dependent diffusion model. This 3D prior, alongside several training strategies, prioritizes the geometry consistency but compromises the texture fidelity. We further propose Bootstrapped Score Distillation to specifically boost the texture. We train a personalized diffusion model, Dreambooth, on the augmented renderings of the scene, imbuing it with 3D knowledge of the scene being optimized. The score distillation from this 3D-aware diffusion prior provides view-consistent guidance for the scene. Notably, through an alternating optimization of the diffusion prior and 3D scene representation, we achieve mutually reinforcing improvements: the optimized 3D scene aids in training the scene-specific diffusion model, which offers increasingly view-consistent guidance for 3D optimization. The optimization is thus bootstrapped and leads to substantial texture boosting. With tailored 3D priors throughout the hierarchical generation, DreamCraft3D generates coherent 3D objects with photorealistic renderings, advancing the state-of-the-art in 3D content generation.

AK

161,530 views • 2 years ago

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 views • 3 years ago

So these researchers figured out you can basically hallucinate 3D cities into existence using just satellite photos & a diffusion model. The problem's pretty straightforward: satellites only see rooftops. Building facades? Invisible. Street-level detail? Doesn't exist. But people want flyable 3D environments, which means you need all that occluded geometry. When I worked on google maps photogrammetry, we could only use satellite-based 3D for isolated stuff like the pyramids - anything city-scale required airplane flyovers. Which is fine until you hit aerial-denied regions where you literally can't fly. Huge chunks of the world just unavailable. Their trick is honestly kind of beautiful. They train gaussian splats on satellite views, but as it descends toward ground level, the renders turn to absolute garbage - artifacts everywhere. Instead of fighting this, they just treat those nightmare renders as the input to a diffusion model. Basically - "hey FLUX, fix this mess." Then here's where it gets clever: they generate multiple diffusion samples per view instead of committing to one. Because any single denoising path is probably wrong in 3D space, but if you generate a couple and let the GS optimization find consensus across them, you get actual geometric consistency. They do this in episodes, curriculum style - start high, gradually descend (hence the name Skyfall-GS!). With each iteration the ground-level views get less fucked. By the end you've got real-time flyable cities that look surprisingly real, and the geometry still matches the satellite input. No 3D training data. No street-level photos. Just satellites + diffusion doing what it does best - filling in the blanks. It's like neural scene completion but actually practical, and it unlocks basically the entire world.

Bilawal Sidhu

242,039 views • 11 months ago

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

177,073 views • 9 months ago

Former Meta Chief AI Scientist Yann LeCun on the three paradigms of machine learning — and why the third is what made ChatGPT possible: Here's each one, and where it breaks. First, supervised learning. You tell the machine the answer. "You show it a picture, let's say of a table, and you tell it this is a table. So it's supervised because you tell it what the correct answer is." Get it wrong, and the machine rewrites itself: "The system computes its output, and if it says something else than table, then it's going to adjust its parameters, its internal structure, so that the output it produces gets closer to the output you want." Repeat at scale and something more than memorisation appears: "Eventually the system will find a way to recognize every image you trained it on, but also images it's never seen that are similar to the one you train it on. This is called a generalization ability." The limit: a human has to supply every single answer. That doesn't scale to the size of the internet. Second, reinforcement learning. You don't give the answer, only a verdict. "You don't tell the system what the correct answer is. You only tell it whether the answer it produced was good or bad." Learning to ride a bike, essentially: "You try to ride a bike and you don't know how to ride the bike and after a while you fall. So you know you did something bad and so you change your strategy a little bit. And eventually you learn how to ride a bike." For years the field assumed this was the closest thing to how animals actually learn. Yann LeCun's verdict: "Now it turns out reinforcement learning is extremely inefficient." It dominates wherever failure is free: "It works really well if you want to train a system to play chess or play go or poker, because you can have the system play millions and millions of games against itself and basically fine-tune itself. But it doesn't really work in the real world." The limit, in one image: "If you want to train a car to drive itself, you're not going to do it with reinforcement learning. It's going to crash thousands of times." On robotics he's careful rather than dismissive: "Reinforcement learning can be part of the solution, but it's not the complete answer. It's not sufficient." Third, self-supervised learning. You tell the machine nothing at all. "And this is what has enabled the recent progress in natural language understanding and chatbots." The strange part is that you stop asking for a task: "You don't train the system to accomplish any particular task. You just train it to basically capture the structure..." The method is deliberate sabotage: "You take a piece of text, you corrupt it in some way, by for example removing some words, and then you train a big neural net to predict the words that are missing." And one narrow version of that trick runs every chatbot on Earth: "A special case of this is that you take a piece of text and the last word in that text is not visible, and so you train the system to predict the last word in that text — and this is the way large language models are trained on." So why did the third one win? Supervised learning needs a human. Reinforcement learning needs a crash. Self-supervised learning needs neither — because the missing word and the correct answer are the same thing. The data grades itself.

Big Brain AI

49,293 views • 1 month ago

Here's a devlog made by an anonymous Chinese fan replicating the surprisingly brand new technique that I developed for detecting asteroids which wound up being so powerful that it can easily track Stealth Fighters from over 100km away even when it’s only using three $30 webcams as sensors meaning it easily outperforms all modern stealth tracking techniques in precision, range and cost. And while this demo is using optical light, this same technique which I call pixel motion to voxel projection, can be used interchangeably with thermal infrared cameras to work at night and also majorly boosts the effectiveness of radar allowing you to track fighters much more effectively through clouds and over the horizon. This technique will also always eventually give the exact location of the target even if the image is blurry as those blurs will always average out from the different perspectives into revealing the precise location of the target in the voxel grid. There is definitely a Mandela effect with this technique as it feels as though it should already exist, especially because at first as it sounds like it is performing triangulation (which has existed for years and is what we do for mocap and tennis ball tracking). But triangulation is entirely separate to this as triangulations only works if you have already identified where the ball is in a 2D image because you’re able to rely on being able to use at least 2 separate high quality cameras which are much closer to the ball making the ball’s apparent size much much bigger and therefore gives you hundreds of pixels to work with which makes it much easier to use object recognition techniques to recognize where it is in the image aka in 2D and then you’re just using the other cameras view to project out lines which intersect in 3D to find out where the ball is in 3D. The major difference is that pixel motion to voxel projection allows you to find where the object is in 3D without having already found it in 2D which is an unbelievable difference as it allows you to use much lower quality cameras together to accumulate data together into 3D space. If this seem like it doesn’t mean much then what it actually means is that you don’t understand what I’m saying as what I’m saying means a LOT in practical terms as it means you go from having to use an imaging system that has to be able to image the object to the point that it is over a hundred total pixels in surface area to have enough data to recognize it to instead be able to use something that is only images the object to be 1 pixel in surface area and only changes the brightness value by 1 value every now and then. I’d recommend an amazing video by DST studios called “Lowlight cameras can’t defeat stealth” if you want a great video which goes over the difficulty of even using telescopes to recognize stealth fighters and why this is so impressive compared to other techniques and ironically it is what inspired me to realize the asteroid tracker I was working on actually could do this. Which brings me to the point that if this wasn’t a new technique then not only would there be at least one example of an asteroid survey that points distant telescopes at the same place at the same time in order to be able to add the light together to detect asteroids which as I was shocked to learn isn’t a thing despite the fact that it would make detecting asteroids trivial by comparison to modern 2D imaging while also having no impact on the normal scientific operations of those surveys other than small changes to scheduling. But there would also be an example of a drone tracker that uses this instead of using the aforementioned high quality zoomable telescope which has to be able to zoom in close enough to be able to recognize a drone. If you want to tell me that this is something that already exists give me an exact example of a product that uses it, not the general outline of a concept that you think it is, the actual product and then also tell me the asteroid survey that uses distant telescopes that point at the exact same place at the exact same time because I can guarantee that if you google what you think uses this you won’t even find the steps of subtracting the images from each other to get motion and will definitely not get the added step of projecting that motion into a voxel grid (It would blow your mind if you found out how Xbox kinect cameras work.) Also I want to make it clear, I’m not saying you should just use web cams to do this, I’m just using them as an example to show you the power of this in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar. Pretty much all of the problems you could think of for this are incredibly easy to overcome if you apply even a small amount of brainpower into fixing the problem. And yes, this gives you the exact location down to the meter of whatever you are tracking even if the image is blurry as those blurs will always average out to the exact location down to the meter in the voxel grid. Which is what makes this technique so powerful since the cost of adding each camera to The network grows linearly while the rate at which each camera gives more information grows exponentially due to the increasing unlikeliness of all of them having more movement in the same place. And given the size of the cameras it really wouldn’t be that hard to hide and network these cameras together in other countries and on sea buoys to know where planes are everywhere in the world. Which brings me to the point that I personally really don’t care about the military uses of this technology, if all it could do is precisely track stealth fighters then I wouldn’t have cared enough to work on it, I could have used any of the many other life saving techniques as the subject of the video, stealth fighters just sounds the most clickable and the scale of the problem is more intuitive to most people and if I did use any of those as subjects for the demo it would inevitably result in the stealth fighter technique being figured out anyway and all of the other uses are so useful that I don't think anyone would reasonably complain about the upside. The real purpose of this video is that since this is a new technique that hasn’t been used to detect stealth fighters despite the billions we have spent on that, then what else can you apply this to that could go on to improve billions of people’s lives that you or others are working on. For example this also allows you to majorly improve the effectiveness of cryo electron microscopy and CT scanners. This part also is kind of hard to explain as it also sounds like it exists but again, when you look through all of the places where you think it is being used you will find that it wasn’t. What I’m saying here isn’t that this is a Radon transform or gaussian splat or whatever, I’m saying that this is able to get new information that wasn’t being accessed before due to the added information about depth you get from the correlation of movement between each perspective which adds to the information that you already have. This allows you to directly subtract foreground and background objects as well as noise faster than you would be able to before and works better than super resolution for your images since super resolution won’t remove foreground and background objects like this does and instead just scales up target, foreground and background objects indiscriminately. And while with enough data Radon transforms or other scanning techniques would eventually get you a correct answer this will get you there a lot faster since those are mostly averaging techniques which average out noise whereas this gets you the ability to directly subtract noise. I’m not expecting you to think that this would do anything but if you try it for yourself you will find that it does majorly improve your ability to perform 3d scans. Again, cryo EM is a field where you would expect this technique to exist but when you look through all the papers on the topic there is no mention of tilting the grid slightly in order to be able to change your perspective slightly on the order of the feature size (if you tilt the grid then you only need precision on the order of an arc minute to do this) and doing multiple exposures from multiple different known tilts and then using those difference images to correlate depth from motion. In fact, in cryo EM you would normally want to do the opposite of this and have your exposures all taken from the same grid angle and just use the variations in how many of the same proteins are oriented in order to be able to scan them for a 3D model but this will generate you far more data faster. There is so much information that I can’t really explain in text so if you have any questions such as why this hasn’t been made before then they will most likely be answered in the video I originally posted which I have added to the end of the first Devlog for your convenience. And again, pretty much all of the problems with the technique can be fixed with a little bit of brainpower, in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar.

ConsistentlyInconsistent

50,925 views • 1 year ago