正在加载视频...
视频加载失败
Researchers at Netflix just released a new AI model It erases objects from video, then rewrites the physics of the entire scene as if that object never existed It's called VOID (Video Object and Interaction Deletion) Current inpainting tools simply paint over the gap left by an object's removal,... show more
47,199 次观看 • 5 个月前 •via X (Twitter)
30 条评论

Github link:

The interesting part isn't the editing, it's what it does to video as evidence. We've spent a century treating footage as the most trustworthy form of testimony. VOID and tools like it quietly retire that assumption. Insurance, journalism, and courts haven't caught up to what that means yet.

Fascinating. VOID goes beyond simple inpainting. It’s actually reasoning about cause and effect in video. Big potential for content editing.

VOID is wild — it doesn’t just paint over the gap, it actually rewrites the physics like the object was never there. This is going to eat a ton of rote VFX cleanup work, but the artists who master it (and the creative decisions around it) are about to level up hard. Future of film looks collaborative, not replaced.

Netflix has the library and the GPU budget. VOID still produces 5 seconds of rough footage. Every physics inpainting demo looks like magic until someone removes a moving car at an intersection.

The real unlock isn't the deletion, it's that it models consequences. Most tools optimize for the immediate problem (fill the hole). VOID optimizes for coherence downstream. What bottleneck in your work assumes you have to choose between speed and logical consistency?

Can you do that for anything Kathleen Kennedy @kennkat, Paramount @paramountplus , et al. have touched and destroyed in the last 20 years - just erase them entirely from our individual and collective consciousness? @TheCriticalDri2

AI is going to dramatically change the film industry. Going to be outputting these movies like we vibe code

日本語でまとめました Summarized in Japanese here:

Physics is the final Frontier of Logic. 🛰️⏳ The Reality Check: Current AI paints over pixels; VOID Refactors the Timeline. In 2026, "Reality" in video isn't a recording—it's a Dynamic Simulation. If you can delete the cause and rewrite the effect, Visual Truth is officially a Legacy Concept. 🦾🗽 The Pulse: Netflix isn't building a tool; they are building the Universal Engine of Alternative Causality. 📈⚡️ Are you watching a "Movie," or are you inside a "Programmable Multiverse"? 👇🚀🏗️

I learned the hard way that "painting over" fails when physics breaks. Real deletion needs circuit breakers—if the scene's state can't reconcile without the object, isolate and rewrite the causal graph instead of hallucinating a patch. Does VOID handle that cascade?

That shift is massive for smarter sports/fitness breakdowns without scene artifacts

rewrites the physics of the entire scene as if that object never existed - that's some next level stuff. in my own experience, when we first added ambient occlusion to our game engine, it changed the way we thought about 3d rendering entirely. the complexity of physics is an even steeper hill to climb.

I run interviews on @TwoSetAI, removing distracting background objects without the weird artifacts current tools leave could be very useful. Any idea if Netflix is releasing this or keeping it internal?

Scoring a 65 percent preference over established tools like Runway shows that users value physical consistency over just visual polish. Even if the output is still a bit rough, the logical consistency of the movement makes it feel much more realistic to the human eye.

This is wild! The idea that VOID can rewrite physics in real-time feels like something out of sci-fi. Can’t wait to see how this evolves with Netflix’s library. Imagine the storytelling possibilities!

wait so it recalculates the shadows and reflections too? thats not inpainting thats basically resimulating the scene. wonder how it handles stuff like a glass on a table where half the refraction paths just vanish

This is the right problem to solve. Erasing the car is trivial. Making the other car keep driving like the crash never happened, that's the hard part nobody talks about

VOID is fascinating. The physics rewrite part is what sets it apart from existing inpainting. Curious if it handles real-time video or only post-production. Either way, this could cut video editing costs dramatically.

Would you trust AI to rewrite reality in your videos, or is that too creepy?

Netflix built a model that erases objects from reality. The object never existed. The physics agree. The implications were not addressed in the press release. 📡

We’re moving from "editing pixels" to "editing reality." For indie creators, this is huge. It basically eliminates the need for expensive reshoots just because a background object ruined a 10/10 take. Rough now, but the training data Netflix has is an unfair advantage

So can Netflix also change the race of an actor or actress in an existing movie?

@grok how does VOID compare to SAM from meta

the jump from “fill the gap with plausible pixels” to “simulate what the scene would look like if this object was never there” is the same jump self-driving made from lane detection to scene prediction. different problem entirely

From experience, this is a huge step. Most tools fake it… this actually understands the physics behind a scene.

Temporal coherence across frames is the hard part — inpainting a single image is solved, but rewriting physics through a whole scene without ghosting is genuinely new. Netflix sitting on this much video data to train on gives them an edge no startup can replicate.

Rowan Cheung highlights Netflix researchers' open-sourced VOID model, which removes objects from videos while simulating realistic physics and causal interactions afterward, unlike basic inpainting that ignores downstream effects like continued motion or falling items. The demo video shows examples such as erasing one car from a collision scene so the other proceeds unimpeded, removing hands from a toy setup to let objects drop naturally, and adjusting pillow stacks after deleting a kettlebell. Built on a 5B-parameter video diffusion model with quadmask conditioning for affected regions, VOID is still limited to short clips but benefits from Netflix's vast training data, positioning it to advance VFX, film post-production, and content editing.

not just erasing pixels but rewriting causality. the model has to understand physics well enough to simulate what would have happened without the object. that's a genuinely hard problem

basically like BTTF1

