Loading video...

Video Failed to Load

Go Home

My biggest Gaussian Splatting scene yet, stitching multiple scans together in one Unity scene with Aras Pranckevičius 🇺🇦🇱🇹 project. Inspired by Block-NeRF! I'm fading between each scan here, but working on seamless transitions next. Kew Palace captured with Insta360 RS 1”. #gaussiansplatting

47,490 views • 2 years ago •via X (Twitter)

10 Comments

grade eterna's profile picture
grade eterna2 years ago

4K YouTube link:

Aras Pranckevičius 🇺🇦🇱🇹's profile picture
Aras Pranckevičius 🇺🇦🇱🇹2 years ago

Whoa. How many millions of splats is this total? (And what’s the performance)

grade eterna's profile picture
grade eterna2 years ago

About 8 mil splats here in total, but only 2 mil at the same time. Deleted a lot of junk with your and @hybridherbst great editing tools! I used the "very high" preset, and it's running about 150fps in HDRP on my 3090.

Mars (parody)'s profile picture
Mars (parody)2 years ago

@aras_p @insta360 I thought splats were already good for streaming, can we not stream segments between scenes anyways?

softyoda's profile picture
softyoda2 years ago

@aras_p @insta360 We need a new and that could render scene by octree (to be dependent of density and render with LOD/HLOD) for massive dataset (10k<200k images) ! Amazing work ! Did you use argument: --position_lr_init 0.000016 --scaling_lr 0.001"

grade eterna's profile picture
grade eterna2 years ago

@aras_p @insta360 Thanks! I used close to default training settings with smaller scan chunks. I haven't had much luck using those scaling arguments with big datasets.

Vincent Bui's profile picture
Vincent Bui2 years ago

@aras_p @insta360 Impressive! It seems the performance is here, are theses files big?

grade eterna's profile picture
grade eterna2 years ago

@aras_p @insta360 Thanks! I used the highest quality preset, so all these scans are around 2GB. But if you used the "very low" preset it would be 18x smaller, and only 110mb!

CorballyGames 🎮 🇮🇪's profile picture
CorballyGames 🎮 🇮🇪2 years ago

@guycalledfrank @aras_p @insta360 I thought this was a drone recording 😳

NeubertIM 📯's profile picture
NeubertIM 📯2 years ago

@aras_p @insta360 🤩🤯

Related Videos

3D scanning and rendering is moving so fast - got my splats up and running and I'm mind blown getting ~100fps for this complex 3D scene ⬇️ 🤯 1. WAY faster than NeRF: For comparison, NeRFs would takes around 10 seconds per frame (!) Instead I'm zipping around with FPV controls without breaking a sweat - though I do crash a few times towards the end of the video lol 2. Old Meets New: Gaussian Splatting is cool in that it fuses classical graphics and deep learning techniques. Like NeRFs, this is still a radiance field - just without the slower (ne)ural rendering part. 3. Explicit Representation: Instead you represent a 3D scene as a collection of ellipsoidal "splats" called gaussians. Each gaussian has a position, size, and color. Rendering in real-time is done by projecting into the image plane and alpha blending. 4. Photorealistic Effects: Gaussian splatting use spherical harmonics to represent the view-dependent effects and lighting - allowing surfaces to change color when viewed from different angles, enabling greater photorealism. It doesn't use a neural network, but the training loop is similar to deep learning. 5. Enables Direct Editing: But it's not just speed - with Gaussian Splatting you also get 3D editing support! So you can select, move, and delete stuff, even relight stuff. This type of editing has been more tedious to do with NeRFs and their implicit black box representations. 📲 More tests cooking! Much more to unpack here including simpler explanations. If you enjoyed this post, you might enjoy my feed: Bilawal Sidhu

Bilawal Sidhu

337,090 views • 2 years ago

Fable 5 and GPT-5.6 built the same scroll-animated website from 1 skill in 32 minutes and only 1 of them made it feel like a film. Same prompt: boutique Japan travel brand, origami style. A subway pulls in, a paper house unfolds into a hotel, a bird takes flight as you scroll. Doing this by hand is brutal multiple videos, matched starting frames, frames ripped out 1 by 1 and synced to scroll position. The skill does all of it: Generate 1 anchor image and approve it every scene inherits the style Turn it into video through the Higgs Field MCP (Seedance), straight from the terminal FFmpeg pulls every frame Each frame maps to scroll position, so your scrollbar becomes the playhead Setup is just connecting the MCP and loading the skill works in Claude Code and Codex. It interviews you first about scenes, budget and mobile, then shows the anchor image before burning a single credit on video. Budget reality: 6 scenes runs ~800 credits and is overkill, 4 scenes is the sweet spot, crop-safe mobile is the cheap path. The verdict came down to transitions. GPT-5.6 Soul built strong scenes with incredible detail inside each one then hard cuts between scene 1 and 2, and again between 2 and 3. Fable 5 stitched them together: wires push out of the top of frame while the next scene blurs in behind, gains depth and locks into place. Same skill, same prompts, 10 generations each. Credit to Peter Wang for the original Scroll World skill open source, now forked with budget tiers and mobile fixes. Soul built 4 scenes. Fable built 1 film.

Spike 1%

44,189 views • 16 days ago

#JANHAE #JanJingjing #EnemiesWithBenefits 💬: which scene do you like the most? 🦊: for me, in the past 3 episodes… let me think first. honestly, i like a lot of scenes. it was fun while filming! i like the scene where they buy stickers for each other. we filmed that scene during q1, and everyone was still… normally, for qs 1-2, they’re like the experimenting phase. oh we filmed it on q2, or q1? one of the first qs, i don’t remember if it’s 1 or 2. qs 1-2 is the period where you’ve officially started for real and are trying to see whether you need to add or adjust anything. it won’t be perfect yet. it’s more like we use it for experimentation 🦊: but that scene was filmed during q1 or 2. at first, everyone probably thought, “let’s try acting it out first! and if anything doesn’t feel right, we’ll adjust it later.” but it turned out that, when we finished filming that scene, everyone (clapped). because it was so cute!!! they said, “oh so cute! so good!” they weren’t worried anymore about us acting in scenes together, our chemistry, or us pairing up. so jing and i immediately high-fived! 🦊: because when we were filming, i already thought it was very cute. and it was the first scene, you know, where we acted together as a pair in cutesy vibes. because most of my scenes in q1… oh yeah this scene was probably q1, and it was mostly my scenes in the office with the nongs (i.e., sales team). i didn’t really get to act with jing too much. so that was probably our first scene together, or it was the scene where we interacted with each other the most at that time. and everyone said we did well so we’re (clapping) happy~ i liked it, that scene was cute!

²²

35,117 views • 2 months ago

Heat (1995) Dir. Michael Mann Mann on the how he prepared the coffee shop scene. "We did two things: We discussed the scene. Then we did some rehearsals, but I was wary because the entire movie is a dialectic that works backward from its last moment... Both men recognize that their next encounter will mean certain death for one of them. Gaining an edge is why they've chosen to meet. So we read the scene a number of times before shooting—not a lot—just looking at it on the page. I didn't want it memorized. My goal was to get them past the unfamiliarity of it. But of course these two already knew it impeccably. We shot that scene with three cameras, two over-the-shoulders and one profile shot, but I found when editing that every time we cut to the profile, the scene lost its one-on-one intensity. I'll often work with multiple cameras, if they're needed. In this case, I knew ahead of time that Pacino and De Niro were so highly attuned to each other that each take would have its own organic unity. Whatever one said, and the specific way he'd say it, would spark a specific reaction in the other. I needed to shoot in such a way that I could use the same take from both angles. What's in the finished film is almost all of take 11—because that has an entirely different integrity and tonality from takes 10, or 9, or 8. All of this begins and ends with scene analysis. It doesn't matter if it's two people in a room or two opposing forces taking over a street. Action comes from drama, and drama is conflict: What's the conflict?"

Gangster Cinema Central

52,284 views • 5 months ago

MVP of Multiview Video → Camera parameters + 3D keypoints. Visualized with Rerun The basic pipeline as of right now looks like this: 1. Capture 🔴 – Using 4 iPhones and an Insta360 Go. iPhone videos are captured via Final Cut Pro Multicam for easy sync and the exocentric view; the Insta360 Go is used for the egocentric view. 2. Sync 🕒 – Custom Gradio app using two Rerun viewers and callbacks for easily aligning frame timestamps so the ego and exo views are aligned. 3. Calibrate 🎯 – Use VGGT from Jianyuan and AI at Meta to get intrinsics/extrinsics for sparse cameras. 4. Estimate 3D 🕺 – Use RTMLib whole‑body keypoint estimator on each frame, then triangulate in 3D. What's missing? 1. No temporal coherence: I’m estimating keypoints one frame at a time and one camera at a time. This leads to a lot of jittering. For now, I plan on adding a One Euro Filter to help with jittering. Long term, I'd want to train a multiview keypoint estimator 2. Kinematic fitting is still missing; this is my next goal. The output will be joint angles, as explored in my previous posts. 3. Missing dense point cloud: VGGT seems to fail for me here. I’m looking to explore using MP‑SFM as a method for generating dense multiview depth maps + normals (plus it has a friendlier license compared to VGGT). 4. Eventually, creation of 4D Gaussian splatting using something akin to DN‑splatter—my long‑term goal is a data engine that provides poses/depths/splats/keypoints/etc.

Pablo Vela

42,785 views • 1 year ago

So how is the Sega Dreamcast's very own native port of Mario Kart 64 coming along? Well, lets take a look at some footage I just captured directly from my DC of jnmartin's latest build! Minor texture corruption is still present on some of the sprites here and there, plus there's no audio yet, BUT MY OH MY! She's running like a dream and looks SHARP as hell at with progressive scan at twice the resolution! You'll immediately notice virtually all of the texture corruption on the text and UI sprites have been fixed. You can now actually see your item box, and clouds are drawn properly in the background, rather than over the top of the scene in the foreground. Somehow, along with fixing all of this in the past few days, jnmartin has also found the time to get the monitors in the backgrounds of Luigi's Raceway and Wario Stadium correctly projecting a view of the scene onto their screens. This somewhat-advanced effect (for the time) was actually done on the N64 (in this game) WITHOUT doing a secondary render-to-texture pass, which requires submitting the scene a second time to the GPU. Instead, the framebuffer was divided into small, tiled chunks (1/6 of its total size), with only a single small chunk getting updated per frame. This way, the amount of direct VRAM access from the CPU (which is typically slow as hell) is kept to a minimum within the duration of a particular frame, with multiple frames being required to update the entire screen. But anyway, she's starting to look and play incredibly well, and those of us with early access to the builds have been thoroughly enjoying revisiting this title on our favorite platform!

Falco Girgis

45,600 views • 1 year ago

Grok Imagine, right now is in my opinion best and fastest ai video generator for the masses. sure, is not perfect, but Rome wasn't built in a day. Maybe ppl from xai or Elon Musk would look on our posts and suggestions for future improvements. What is a must (for advanced users into ai video generation, been doing this game since 2022) .. 1. for longer movies , we need an option to organize like a project style, and to be able to add main prompts like the niche of the current movie, the character description and to be able to select a custom seed so we can have consistency of the characters. 2. we have now 6 seconds generation ( saw Elon promised 15 seconds soon).. BUT when we generate long movies, we end up with lots of scenes... what grok needs for the same project of the movie, would be a First Frame -Last frame scene interpolation between the scenes (take last frame from scene one, and first frame from scene 2 and generate a mid scene that would merge scene 1 with scene 2 .. and continue for the other scenes (this could be very easy implemented with some python lines of code , like before spitting final video, select all scenes.. extract frames etc etc etc etc.. simple af, when u have all scenes + the interpolation scenes combine evrything with ffmpeg ). 3.. list is long... and i dind't finished my coffee yet, so here is a grok TEXT to video short movie (coz lol u hit the limit for today). Prompts i used for each scene are a little more advanced, so i can see what grok is able to do .. the prompts used are like this (can;t post all due to X limits ) : { "scene_1": { "global_cinematography": "Ultra-realistic Hollywood cyberpunk thriller in the vein of The Matrix (1999) and Blade Runner 2049 (2017), shot on Arri Alexa LF with anamorphic lenses for widescreen 2.39:1 aspect ratio, 24fps for fluid motion, desaturated palette dominated by cool blues, greens, and high-contrast neon reds piercing perpetual smog-choked night. Consistent VFX pipeline: Procedural green code cascades, photorealistic cybernetic augmentations with subsurface scattering, physics-based rain and particle simulations. Lighting paradigm: Volumetric god rays through haze, practical lens flares from holograms, rim lighting on metallic surfaces for depth. Sound integration: Pulsing industrial synth score with digital glitches, rain patter syncing to code interference, metallic echoes underscoring dialogue. Transitions: Seamless glitch wipes or matrix symbol dissolves ensuring narrative continuity, each scene's final beat priming the next for unbroken tension flow. Continuity directive: Scenes chain via lingering elements—rain droplets from prior shots persisting, Nova's silhouette echoing across cuts, HUD overlays threading flashbacks to present, escalating glitch distortions building to climax rupture—maintaining spatial and temporal cohesion in Neo-Tokyo's underbelly.", "shot": { "composition": "Wide aerial drone shot with 35mm wide-angle anamorphic lens on Arri Alexa LF, high dynamic range capturing smog gradients and rain refraction for immersive dystopian establishment, foreground skyscraper edges framing the descent path", "camera_motion": "Controlled descending tilt-push through layered haze, subtle forward momentum building velocity into street-level convergence, priming alley reveal for Scene 2 silhouette emergence" }, "subject": { "description": "Neo-Tokyo's jagged circuit-board skyscrapers thrusting into smog-veiled void, rain-lashed surfaces mirroring erratic neon pulses; faint pedestrian phantoms below as harbingers of oblivious simulation", "wardrobe": "null" }, "scene": { "location": "Shadowed aerial vantage over Neo-Tokyo underbelly, continuity hook from global haze motif", "time_of_day": "Perpetual neon-twilight under storm overcast, syncing with all scenes' eternal dusk", "environment": "Thick smog banks parting reluctantly, acid rain sheets cascading in synchronized sheets with volumetric depth, holographic billboards stuttering in the distance to echo Scene 7 flicker" }, "visual_details": { "action": "Drone pierces urban canopy, unveiling rain-assaulted sprawl where neon bleeds into puddles like corrupted signals, distant alley haze teasing Nova's imminent step-forward in Scene 2", "props": "Circuit-etched tower facades with embedded LED veins flickering erratically, overflowing industrial gutters spewing iridescent chemical runoff, wind-scattered debris hinting at skirmish aftermath", "action_sequence": [ {"0-1s": "High hover frames smog-piercing spires, rain droplets streak lens in slow-mo refraction"}, {"1-2s": "Descent accelerates, haze thins to reveal neon-veined edges glowing faintly blue"}, {"2-3s": "Tilt reveals grid below, rooftops hammered in static-burst impacts syncing to score pulse"}, {"3-4s": "Forward push threads alley corridors, Mandarin signs initial flicker priming Scene 7"}, {"4-5s": "Pedestrians sharpen as wireframe ghosts, AR visors glinting obliviously"}, {"5-6s": "Level to ground haze, Nova's trench silhouette materializes at frame's vanishing point, coat billow lingering into Scene 2 track"} ] }, "cinematography": { "lighting": "Desaturated neon primaries with volumetric god rays slicing haze for ethereal isolation, rain speculars adding dynamic highlights consistent across wet surfaces", "tone": "Oppressive immersion yielding to rebellious spark—global cyber-noir dread laced with glitch anticipation, flowing seamlessly to Nova's personal emergence" } }, "scene_2": { "global_cinematography": "Ultra-realistic Hollywood cyberpunk thriller in the vein of The Matrix (1999) and Blade Runner 2049 (2017), shot on Arri Alexa LF with anamorphic lenses for widescreen 2.39:1 aspect ratio, 24fps for fluid motion, desaturated palette dominated by cool blues, greens, and high-contrast neon reds piercing perpetual smog-choked night. Consistent VFX pipeline: Procedural green code cascades, photorealistic cybernetic augmentations with subsurface scattering, physics-based rain and particle simulations. Lighting paradigm: Volumetric god rays through haze, practical lens flares from holograms, rim lighting on metallic surfaces for depth. Sound integration: Pulsing industrial synth score with digital glitches, rain patter syncing to code interference, metallic echoes underscoring dialogue. Transitions: Seamless glitch wipes or matrix symbol dissolves ensuring narrative continuity, each scene's final beat priming the next for unbroken tension flow. Continuity directive: Scenes chain via lingering elements—rain droplets from prior shots persisting, Nova's silhouette echoing across cuts, HUD overlays threading flashbacks to present, escalating glitch distortions building to climax rupture—maintaining spatial and temporal cohesion in Neo-Tokyo's underbelly.", "shot": { "composition": "Low-angle tracking push with 50mm anamorphic prime on Arri Alexa LF, heroic distortion compressing background alley into claustrophobic funnel, foreground rain blur veiling initial fog for continuity from Scene 1 descent", "camera_motion": "Fluid forward Steadicam arc from lingering Scene 1 haze, subtle left profile tilt to frame Nova against graffiti wall, pulling back slightly to hold environmental depth into Scene 3 orbit" }, "subject": { "description": "Nova, 30s hybrid rebel with scarred synthetic pallor, cropped black hair rain-matted, holographic irises scanning with latent data flickers; sleek titanium limbs rune-etched in dormant blue", "wardrobe": "Sodden black trench coat with frayed hems from Scene 1 debris scatter, high collar shadowing jawline for motif continuity" }, "scene": { "location": "Graffiti-choked alley continuation from Scene 1 street convergence, Neo-Tokyo underbelly", "time_of_day": "Eternal neon-dusk syncing global palette", "environment": "Fog banks rolling from industrial vents as Scene 1 smog extension, wet cobblestones rippling with residual aerial rain patterns" }, "visual_details": { "action": "Nova materializes from Scene 1's terminal haze, striding assertively into sodium glow with metallic glint, coat hem dragging puddles to splash forward—teasing Scene 3 facial trace", "props": "Luminescent 'GLITCH THE SYSTEM' graffiti echoing from Scene 1 signs, overhead hover-traffic hum persisting from aerial hum", "action_sequence": [ {"0-1s": "Fog swirl from Scene 1 yields Nova's silhouette, boot first impacting puddle"}, {"1-2s": "Full stride forward, coat hem trails iridescent wake linking to blood drip in Scene 9"}, {"2-3s": "Titanium forearm catches neon, runes sequential-pulse awakening blue continuity"}, {"3-4s": "Holographic eyes iris-scan, reflecting alley code fragments priming Scene 4 overlay"}, {"4-5s": "Rain beads contour synthetic skin, parting at seams for Scene 3 macro journey"}, {"5-6s": "Profile lean against wall, vapor breath hangs, posture straightening into Scene 5 OTS"} ] }, "cinematography": { "lighting": "Harsh sodium sidelight rimming form per global motif, cool rune fill softening human remnants, prismatic rain refractions tying to Scene 1 aerial streaks", "tone": "Defiant grace in simulated decay—cyber-noir intimacy building personal stakes, camera arc ensuring spatial flow to close-up revelation" } }, etc etc etc up to scene 16. you got the point

NFK

3,351,468 views • 8 months ago

"Hah - generative ai can't even make an image of a hand with the right number of fingers.." "Stop pushing this slop" It's way past the point now where it must be clear to everyone, that generative ai is here to stay AND that the quality will continue to increase. I've been talking about this trajectory for years now, and I've been working towards finding ways to combine the strength of these models, with the best of what I love about "old-school" creation. Building with my hands, moving a pencil across the paper and seeing shapes emerge, moving a building slightly to the right to get just that composition I had in mind. Being fully immersed in a scene I'm building in VR, being inspired by the immersion to take the story in a new direction. For years I've been talking about how powerful the combination of 3d and generative ai is, be it traditional 3d, SDF volumes in Dreams or Gaussian splats - with experiments around using V2V as a "render pass" or with experiments around realtime ai. Enough talk you might think, where's the proof? It's all around us these days honestly and here's a small test I did during some OOO. Blender MPC + Fable - a pretty powerful combination! With a bit of Google Omni Fast on top as a "render" pass. What do you think of where this is heading? Hopeful, disheartened, inspired or the opposite? Can you imagine working with tools like this in a way where we still retain the human "spark" and the creative nerve that makes each persons creation unique?

Martin Nebelong

48,179 views • 7 days ago

Fast Company just published a great piece on World Labs , Fei-Fei Li , Marble, and the idea that spatial intelligence / world models may be one of the next big shifts in AI. I was happy to be quoted in the article, but I also wanted to share more context about my own experience with World Labs and Marble, and why this direction is especially interesting to me. My starting point: volumetric capture — For the past few years I’ve been exploring and using volumetric capture and reconstruction (photogrammetry, NeRFs, 3D Gaussian Splats) mostly capturing locations around Montreal. Alleys, museums, urban interiors. I love every step of it: the capture itself, the pipeline, and what can be done with the output. Turning real spaces into real-time explorable systems. I do this personally, sharing explorations here, and professionally as chief technologist, and co-founder of Dpt. Physical reality + generative manipulation — In my work I’m especially drawn to mixing physical reality with generative and digital manipulation: using physical interfaces (light, clay, ink, ... ) to drive generative AI pipelines, building mixed reality prototypes that reshape your surroundings, or starting from real captured spaces and transforming them using tools like Marble. Like many people, I saw the World Labs announcement on Twitter in September 2024, and Marble when it surfaced in early December. But by then, I already had a sense something was coming. The first conversation — As someone deep into volumetric capture and radiance fields, I obviously knew about Ben Mildenhall and his pioneering work on NeRF. To my surprise, Ben reached out to me in late June 2024. He’d been following some of my experiments and wanted to chat about my process and workflows and how I was using this “stuff” creatively. At that point he didn’t share what he was building, but we had a genuinely great conversation about radiance fields, AI, and my work. He was curious about the creative perspective, not just the technical one. When the World Labs announcement dropped a few months later, it all made sense. I understood what Ben had been working on, and why the creative angle mattered to them. Then in August 2025, he invited me to try the Marble beta, and I’ve been experimenting with it since. Experimenting with Marble — The first thing I used Marble for was materializing scene and world concepts during ideation at the studio, and seeing if and how it could fit into our production pipeline. In parallel, I dove into a series of experiments focused on world manipulation: starting from real captured spaces and transforming them using Marble. I’d already been exploring that idea using img2img diffusion with ControlNet on NeRF renders, real-time video streams, and even mixed reality using headset camera feeds. But Marble brings something different. It generates persistent, spatially cohesive 3D worlds that can be rendered in real time across a wide range of devices. That’s a real shift. Experiment 01: Parallel Realities — The first experiment, Parallel Realities, starts from a volumetric capture of a real location, reconstructed as 3D Gaussian Splats. Using Marble, I generate an alternate version of that same space, something informed by the original architecture: abandoned, nature-reclaimed, alternate era. Then, using Spark (World Labs’ 3D Gaussian Splatting renderer for THREE.js) I make both realities coexist in the same spatial coordinate system. From there, I use a portal UX mechanic to let the user step between the real reconstruction and the Marble-generated version. Experiment 02: Hidden Depth The second experiment, Hidden Depth, does not transform a space as much as expand it. A captured location has a visual boundary (a mural, a doorway, a dark corridor) and Marble generates what exists beyond it. For example: a Montreal alley has a painted mural; step through it and you’re inside a world informed by what is actually depicted there. World Labs showcased part of this work here: And in their Spark 2.0 post: The project page is here: Why this matters to me — Being able to start from a real 3D Gaussian Splat scene and manipulate it with Marble opens up a lot of ideas. The 3DGS pipeline is becoming an increasingly compelling foundation for exploration, experimentation, and storytelling. What matters most to me right now is more control. The more I can steer the generated scene or world, the more useful the tool becomes. I want more features like the already existing multiple input images and Chisel, the blockout-based approach. I would like better local control, the ability to expand a generated world more and more while preserving coherence, and the ability to directly import 3D Gaussian Splat scenes to be used as a starting point. I want more ways to shape the result, not just a “prompt and hope” approach. — It is exciting to see this field moving from research and demos toward actual creative workflows.

Hugues Bruyère

69,337 views • 1 month ago