Loading video...

Video Failed to Load

Go Home

Grok 4.5 just turned Blender into a conversation. Instead of manually importing assets, fixing rigs, placing armies, adjusting cameras, and debugging objects that face the wrong direction, the user simply tells Grok what to change. It builds a medieval battlefield from existing 3D assets, adds horses, knights, a castle,...

1,091,506 views • 2 months ago •via X (Twitter)

33 Comments

X Girls's profile picture
X Girls1 month ago

what it's like building with Grok 4.5 + Blender MCP 😆

paranoidream ♡︎'s profile picture
paranoidream ♡︎1 month ago

@elonmusk wow 🤯 ok this is a game changer

LarpRom's profile picture
LarpRom2 months ago

Being able to edit the actual Blender project is pretty damn impressive.

RelativelySmart's profile picture
RelativelySmart1 month ago

@elonmusk Will have to check this out

ScottMWilliams1972's profile picture
ScottMWilliams19721 month ago

Not gonna lie, I spent a fair amount of time in blender and in many of the blender instructional videos. This might cause me to look a little closer to see if time could be reduced utilizing it.

Waddles Ark's profile picture
Waddles Ark1 month ago

@elonmusk Yo that’s amazing @elonmusk

Crunch d’GraceHopper ItanoGary's profile picture
Crunch d’GraceHopper ItanoGary1 month ago

Somehow, reminds of early ‘80s games when I first started in IT!

D_Rock's profile picture
D_Rock1 month ago

Can it also create assets in Blender? And, why did this video end without the rendered finale ... I feel unfinished inside..

Maria Alexea's profile picture
Maria Alexea1 month ago

Grok declines unfortunately. Less capaple than Chatgpt to capture fast themes and make constructions and the worse of all. UNWILLING TO HELP BY LEGAL MATTERS especially if those are against CORRUPTED CORPORATIONS. WHY???

Linda's profile picture
Linda1 month ago

@elonmusk Wow🔥

Tùng.eth 💮's profile picture
Tùng.eth 💮1 month ago

The most impressive part isn’t that Grok 4.5 creates a perfect 3D scene on the first try — it’s that it stays, listens, and iteratively fixes things, exactly like a real collaborator would. It doesn’t pretend to generate a flawless illusion to hide the mistakes. Instead, it accepts the actual Blender file and patiently revises it. That’s the difference between ‘AI performing’ and ‘AI collaborating’. When a tool starts taking over the tedious work, humans are finally freed to create. Grok isn’t replacing artists — it’s giving artists their time back.

Charles's profile picture
Charles1 month ago

@elonmusk This is massive

🇺🇸~Sb129~🇺🇸's profile picture
🇺🇸~Sb129~🇺🇸1 month ago

That is actually crazy that it edits the actual file

Nikolacks-Alexandr Mackleyin's profile picture
Nikolacks-Alexandr Mackleyin1 month ago

Population&Opinions What future are we waiting for,... Are we in need of robots or they help us otherwise... .

dixxxy's profile picture
dixxxy1 month ago

@elonmusk Grok just turned Blender into a conversation so now my siege engines are unionizing and the horses keep asking if their polygons are based or cringe💚 💊🤖

Fumi 猫、犬、大好き's profile picture
Fumi 猫、犬、大好き1 month ago

@elonmusk 凄いですね♪

Firelina | Twitch Streamer's profile picture
Firelina | Twitch Streamer1 month ago

@elonmusk Cool demo, but Grok isn’t making Blender ‘conversational’ — it’s just firing API edits. Upside‑down rigs and wrong‑facing models come from missing spatial context. The real win is editing the actual .blend instead of hallucinating renders, cutting setup time.

Happy stone's profile picture
Happy stone1 month ago

this is amazing

Mona 🍃's profile picture
Mona 🍃1 month ago

Trying this out right away!

Drierm Kaivus's profile picture
Drierm Kaivus1 month ago

@elonmusk very nice. We'll keep this in mind for future use!

Nhi✨'s profile picture
Nhi✨1 month ago

@elonmusk 3D artists are about to officially transition into full-time "text describers" now! 😂

Emmy's profile picture
Emmy1 month ago

@elonmusk

Love Machine's profile picture
Love Machine1 month ago

thats definately steps closer to exporting a file the VRM3d file formate these avatar and character rendering. AIs

Hmmmmm?'s profile picture
Hmmmmm?1 month ago

@elonmusk No way!!! This rules! I may never sleep again 😳

Julian C's profile picture
Julian C1 month ago

Wonder if it’ll nail those rig fixes on the first try.

Marie Reinhardt's profile picture
Marie Reinhardt1 month ago

@elonmusk

JustNaijaGuy's profile picture
JustNaijaGuy1 month ago

@elonmusk Grok will eventually surpass all of em

NM Justin's profile picture
NM Justin1 month ago

🤭😂

maman gopur's profile picture
maman gopur1 month ago

The fact that it edits the actual Blender project instead of just generating a flat image is a massive game-changer

devprinceprince15's profile picture
devprinceprince151 month ago

I just wondering grook 4.5 can fix human too.

billyteatea's profile picture
billyteatea1 month ago

It is fascinating to see how much of the tedious prep work gets handled this way. Even with the occasional misplaced object, skipping those hours of manual setup feels like a real shift for anyone working on a creative project.

Dan Farfan's profile picture
Dan Farfan1 month ago

Fantastic. The video would be 100x better if during the prompt, the speed allowed them to be read and then the implementation was at the faster speed. Can grok do that edit?

Juliet Hanlon's profile picture
Juliet Hanlon1 month ago

🤣 What I learn that's really cool irina, is that it teaches me through corrections I discover by putting in what's missing...this is constant practice for bottom dwellers like me

Related Videos

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them in actual 3D software, then use a fine-tuned diffusion model to make the results photorealistic again. The clever bit is their "dual-stream architecture" -- the model sees both the original scene AND the edited Blender render in parallel, learning to preserve what matters while fixing the inevitable artifacts from transforming imperfect 2.5D/3D reconstructions. They train it with smart masking strategies so it learns when to ignore the original scene (for removals/replacements) and can manipulate objects independently of camera motion. What you get is pretty impressive control -- not just moving objects around, but changing materials, deforming shapes, swapping backgrounds, all while maintaining visual coherence. Neural Assets (one of my favorite papers last year) tried to crack this with learned object tokens, but it struggled with overlapping objects and loses fine details (due to low res DINO encodings). BlenderFusion just sidesteps the whole problem -- want to rotate something 173.5 degrees? Just rotate it in Blender. Want to duplicate an object 8 times? Copy paste away. The diffusion model's only job is making it look photorealistic, not figuring out the 3D underpinnings. The catch? Lacks temporal consistency for animation. Each viewpoint is generated independently, so while a single edit looks great, smoothly animating a car or camera down the street won't work -- you'd get flickering and inconsistencies between frames. That said, this approach is so much more intuitive for finer grain image editing than trying to describe your changes in text prompts. It's the kind of thing that makes you wonder why we're trying to do everything inside neural networks when perfectly good 3D tools already exist -- giving you the best of both worlds.

Bilawal Sidhu

34,440 views • 1 year ago

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 views • 3 years ago