Loading video...

Video Failed to Load

Go Home

My biggest Gaussian Splatting scene yet, stitching multiple scans together in one Unity scene with Aras Pranckevičius 🇺🇦🇱🇹 project. Inspired by Block-NeRF! I'm fading between each scan here, but working on seamless transitions next. Kew Palace captured with Insta360 RS 1”. #gaussiansplatting

47,490 views • 2 years ago •via X (Twitter)

10 Comments

grade eterna's profile picture
grade eterna2 years ago

4K YouTube link:

Aras Pranckevičius 🇺🇦🇱🇹's profile picture
Aras Pranckevičius 🇺🇦🇱🇹2 years ago

Whoa. How many millions of splats is this total? (And what’s the performance)

grade eterna's profile picture
grade eterna2 years ago

About 8 mil splats here in total, but only 2 mil at the same time. Deleted a lot of junk with your and @hybridherbst great editing tools! I used the "very high" preset, and it's running about 150fps in HDRP on my 3090.

Mars (parody)'s profile picture
Mars (parody)2 years ago

@aras_p @insta360 I thought splats were already good for streaming, can we not stream segments between scenes anyways?

softyoda's profile picture
softyoda2 years ago

@aras_p @insta360 We need a new and that could render scene by octree (to be dependent of density and render with LOD/HLOD) for massive dataset (10k<200k images) ! Amazing work ! Did you use argument: --position_lr_init 0.000016 --scaling_lr 0.001"

grade eterna's profile picture
grade eterna2 years ago

@aras_p @insta360 Thanks! I used close to default training settings with smaller scan chunks. I haven't had much luck using those scaling arguments with big datasets.

Vincent Bui's profile picture
Vincent Bui2 years ago

@aras_p @insta360 Impressive! It seems the performance is here, are theses files big?

grade eterna's profile picture
grade eterna2 years ago

@aras_p @insta360 Thanks! I used the highest quality preset, so all these scans are around 2GB. But if you used the "very low" preset it would be 18x smaller, and only 110mb!

CorballyGames 🎮 🇮🇪's profile picture
CorballyGames 🎮 🇮🇪2 years ago

@guycalledfrank @aras_p @insta360 I thought this was a drone recording 😳

NeubertIM 📯's profile picture
NeubertIM 📯2 years ago

@aras_p @insta360 🤩🤯

Related Videos

3D scanning and rendering is moving so fast - got my splats up and running and I'm mind blown getting ~100fps for this complex 3D scene ⬇️ 🤯 1. WAY faster than NeRF: For comparison, NeRFs would takes around 10 seconds per frame (!) Instead I'm zipping around with FPV controls without breaking a sweat - though I do crash a few times towards the end of the video lol 2. Old Meets New: Gaussian Splatting is cool in that it fuses classical graphics and deep learning techniques. Like NeRFs, this is still a radiance field - just without the slower (ne)ural rendering part. 3. Explicit Representation: Instead you represent a 3D scene as a collection of ellipsoidal "splats" called gaussians. Each gaussian has a position, size, and color. Rendering in real-time is done by projecting into the image plane and alpha blending. 4. Photorealistic Effects: Gaussian splatting use spherical harmonics to represent the view-dependent effects and lighting - allowing surfaces to change color when viewed from different angles, enabling greater photorealism. It doesn't use a neural network, but the training loop is similar to deep learning. 5. Enables Direct Editing: But it's not just speed - with Gaussian Splatting you also get 3D editing support! So you can select, move, and delete stuff, even relight stuff. This type of editing has been more tedious to do with NeRFs and their implicit black box representations. 📲 More tests cooking! Much more to unpack here including simpler explanations. If you enjoyed this post, you might enjoy my feed: Bilawal Sidhu

Bilawal Sidhu

337,090 views • 2 years ago

Fable 5 and GPT-5.6 built the same scroll-animated website from 1 skill in 32 minutes and only 1 of them made it feel like a film. Same prompt: boutique Japan travel brand, origami style. A subway pulls in, a paper house unfolds into a hotel, a bird takes flight as you scroll. Doing this by hand is brutal multiple videos, matched starting frames, frames ripped out 1 by 1 and synced to scroll position. The skill does all of it: Generate 1 anchor image and approve it every scene inherits the style Turn it into video through the Higgs Field MCP (Seedance), straight from the terminal FFmpeg pulls every frame Each frame maps to scroll position, so your scrollbar becomes the playhead Setup is just connecting the MCP and loading the skill works in Claude Code and Codex. It interviews you first about scenes, budget and mobile, then shows the anchor image before burning a single credit on video. Budget reality: 6 scenes runs ~800 credits and is overkill, 4 scenes is the sweet spot, crop-safe mobile is the cheap path. The verdict came down to transitions. GPT-5.6 Soul built strong scenes with incredible detail inside each one then hard cuts between scene 1 and 2, and again between 2 and 3. Fable 5 stitched them together: wires push out of the top of frame while the next scene blurs in behind, gains depth and locks into place. Same skill, same prompts, 10 generations each. Credit to Peter Wang for the original Scroll World skill open source, now forked with budget tiers and mobile fixes. Soul built 4 scenes. Fable built 1 film.

Spike 1%

45,052 views • 1 month ago

#JANHAE #JanJingjing #EnemiesWithBenefits 💬: which scene do you like the most? 🦊: for me, in the past 3 episodes… let me think first. honestly, i like a lot of scenes. it was fun while filming! i like the scene where they buy stickers for each other. we filmed that scene during q1, and everyone was still… normally, for qs 1-2, they’re like the experimenting phase. oh we filmed it on q2, or q1? one of the first qs, i don’t remember if it’s 1 or 2. qs 1-2 is the period where you’ve officially started for real and are trying to see whether you need to add or adjust anything. it won’t be perfect yet. it’s more like we use it for experimentation 🦊: but that scene was filmed during q1 or 2. at first, everyone probably thought, “let’s try acting it out first! and if anything doesn’t feel right, we’ll adjust it later.” but it turned out that, when we finished filming that scene, everyone (clapped). because it was so cute!!! they said, “oh so cute! so good!” they weren’t worried anymore about us acting in scenes together, our chemistry, or us pairing up. so jing and i immediately high-fived! 🦊: because when we were filming, i already thought it was very cute. and it was the first scene, you know, where we acted together as a pair in cutesy vibes. because most of my scenes in q1… oh yeah this scene was probably q1, and it was mostly my scenes in the office with the nongs (i.e., sales team). i didn’t really get to act with jing too much. so that was probably our first scene together, or it was the scene where we interacted with each other the most at that time. and everyone said we did well so we’re (clapping) happy~ i liked it, that scene was cute!

²²

35,117 views • 2 months ago

Heat (1995) Dir. Michael Mann Mann on the how he prepared the coffee shop scene. "We did two things: We discussed the scene. Then we did some rehearsals, but I was wary because the entire movie is a dialectic that works backward from its last moment... Both men recognize that their next encounter will mean certain death for one of them. Gaining an edge is why they've chosen to meet. So we read the scene a number of times before shooting—not a lot—just looking at it on the page. I didn't want it memorized. My goal was to get them past the unfamiliarity of it. But of course these two already knew it impeccably. We shot that scene with three cameras, two over-the-shoulders and one profile shot, but I found when editing that every time we cut to the profile, the scene lost its one-on-one intensity. I'll often work with multiple cameras, if they're needed. In this case, I knew ahead of time that Pacino and De Niro were so highly attuned to each other that each take would have its own organic unity. Whatever one said, and the specific way he'd say it, would spark a specific reaction in the other. I needed to shoot in such a way that I could use the same take from both angles. What's in the finished film is almost all of take 11—because that has an entirely different integrity and tonality from takes 10, or 9, or 8. All of this begins and ends with scene analysis. It doesn't matter if it's two people in a room or two opposing forces taking over a street. Action comes from drama, and drama is conflict: What's the conflict?"

Gangster Cinema Central

52,284 views • 5 months ago

MVP of Multiview Video → Camera parameters + 3D keypoints. Visualized with Rerun The basic pipeline as of right now looks like this: 1. Capture 🔴 – Using 4 iPhones and an Insta360 Go. iPhone videos are captured via Final Cut Pro Multicam for easy sync and the exocentric view; the Insta360 Go is used for the egocentric view. 2. Sync 🕒 – Custom Gradio app using two Rerun viewers and callbacks for easily aligning frame timestamps so the ego and exo views are aligned. 3. Calibrate 🎯 – Use VGGT from Jianyuan and AI at Meta to get intrinsics/extrinsics for sparse cameras. 4. Estimate 3D 🕺 – Use RTMLib whole‑body keypoint estimator on each frame, then triangulate in 3D. What's missing? 1. No temporal coherence: I’m estimating keypoints one frame at a time and one camera at a time. This leads to a lot of jittering. For now, I plan on adding a One Euro Filter to help with jittering. Long term, I'd want to train a multiview keypoint estimator 2. Kinematic fitting is still missing; this is my next goal. The output will be joint angles, as explored in my previous posts. 3. Missing dense point cloud: VGGT seems to fail for me here. I’m looking to explore using MP‑SFM as a method for generating dense multiview depth maps + normals (plus it has a friendlier license compared to VGGT). 4. Eventually, creation of 4D Gaussian splatting using something akin to DN‑splatter—my long‑term goal is a data engine that provides poses/depths/splats/keypoints/etc.

Pablo Vela

42,785 views • 1 year ago

So how is the Sega Dreamcast's very own native port of Mario Kart 64 coming along? Well, lets take a look at some footage I just captured directly from my DC of jnmartin's latest build! Minor texture corruption is still present on some of the sprites here and there, plus there's no audio yet, BUT MY OH MY! She's running like a dream and looks SHARP as hell at with progressive scan at twice the resolution! You'll immediately notice virtually all of the texture corruption on the text and UI sprites have been fixed. You can now actually see your item box, and clouds are drawn properly in the background, rather than over the top of the scene in the foreground. Somehow, along with fixing all of this in the past few days, jnmartin has also found the time to get the monitors in the backgrounds of Luigi's Raceway and Wario Stadium correctly projecting a view of the scene onto their screens. This somewhat-advanced effect (for the time) was actually done on the N64 (in this game) WITHOUT doing a secondary render-to-texture pass, which requires submitting the scene a second time to the GPU. Instead, the framebuffer was divided into small, tiled chunks (1/6 of its total size), with only a single small chunk getting updated per frame. This way, the amount of direct VRAM access from the CPU (which is typically slow as hell) is kept to a minimum within the duration of a particular frame, with multiple frames being required to update the entire screen. But anyway, she's starting to look and play incredibly well, and those of us with early access to the builds have been thoroughly enjoying revisiting this title on our favorite platform!

Falco Girgis

45,613 views • 1 year ago

Grok Imagine, right now is in my opinion best and fastest ai video generator for the masses. sure, is not perfect, but Rome wasn't built in a day. Maybe ppl from xai or Elon Musk would look on our posts and suggestions for future improvements. What is a must (for advanced users into ai video generation, been doing this game since 2022) .. 1. for longer movies , we need an option to organize like a project style, and to be able to add main prompts like the niche of the current movie, the character description and to be able to select a custom seed so we can have consistency of the characters. 2. we have now 6 seconds generation ( saw Elon promised 15 seconds soon).. BUT when we generate long movies, we end up with lots of scenes... what grok needs for the same project of the movie, would be a First Frame -Last frame scene interpolation between the scenes (take last frame from scene one, and first frame from scene 2 and generate a mid scene that would merge scene 1 with scene 2 .. and continue for the other scenes (this could be very easy implemented with some python lines of code , like before spitting final video, select all scenes.. extract frames etc etc etc etc.. simple af, when u have all scenes + the interpolation scenes combine evrything with ffmpeg ). 3.. list is long... and i dind't finished my coffee yet, so here is a grok TEXT to video short movie (coz lol u hit the limit for today). Prompts i used for each scene are a little more advanced, so i can see what grok is able to do .. the prompts used are like this (can;t post all due to X limits ) : { "scene_1": { "global_cinematography": "Ultra-realistic Hollywood cyberpunk thriller in the vein of The Matrix (1999) and Blade Runner 2049 (2017), shot on Arri Alexa LF with anamorphic lenses for widescreen 2.39:1 aspect ratio, 24fps for fluid motion, desaturated palette dominated by cool blues, greens, and high-contrast neon reds piercing perpetual smog-choked night. Consistent VFX pipeline: Procedural green code cascades, photorealistic cybernetic augmentations with subsurface scattering, physics-based rain and particle simulations. Lighting paradigm: Volumetric god rays through haze, practical lens flares from holograms, rim lighting on metallic surfaces for depth. Sound integration: Pulsing industrial synth score with digital glitches, rain patter syncing to code interference, metallic echoes underscoring dialogue. Transitions: Seamless glitch wipes or matrix symbol dissolves ensuring narrative continuity, each scene's final beat priming the next for unbroken tension flow. Continuity directive: Scenes chain via lingering elements—rain droplets from prior shots persisting, Nova's silhouette echoing across cuts, HUD overlays threading flashbacks to present, escalating glitch distortions building to climax rupture—maintaining spatial and temporal cohesion in Neo-Tokyo's underbelly.", "shot": { "composition": "Wide aerial drone shot with 35mm wide-angle anamorphic lens on Arri Alexa LF, high dynamic range capturing smog gradients and rain refraction for immersive dystopian establishment, foreground skyscraper edges framing the descent path", "camera_motion": "Controlled descending tilt-push through layered haze, subtle forward momentum building velocity into street-level convergence, priming alley reveal for Scene 2 silhouette emergence" }, "subject": { "description": "Neo-Tokyo's jagged circuit-board skyscrapers thrusting into smog-veiled void, rain-lashed surfaces mirroring erratic neon pulses; faint pedestrian phantoms below as harbingers of oblivious simulation", "wardrobe": "null" }, "scene": { "location": "Shadowed aerial vantage over Neo-Tokyo underbelly, continuity hook from global haze motif", "time_of_day": "Perpetual neon-twilight under storm overcast, syncing with all scenes' eternal dusk", "environment": "Thick smog banks parting reluctantly, acid rain sheets cascading in synchronized sheets with volumetric depth, holographic billboards stuttering in the distance to echo Scene 7 flicker" }, "visual_details": { "action": "Drone pierces urban canopy, unveiling rain-assaulted sprawl where neon bleeds into puddles like corrupted signals, distant alley haze teasing Nova's imminent step-forward in Scene 2", "props": "Circuit-etched tower facades with embedded LED veins flickering erratically, overflowing industrial gutters spewing iridescent chemical runoff, wind-scattered debris hinting at skirmish aftermath", "action_sequence": [ {"0-1s": "High hover frames smog-piercing spires, rain droplets streak lens in slow-mo refraction"}, {"1-2s": "Descent accelerates, haze thins to reveal neon-veined edges glowing faintly blue"}, {"2-3s": "Tilt reveals grid below, rooftops hammered in static-burst impacts syncing to score pulse"}, {"3-4s": "Forward push threads alley corridors, Mandarin signs initial flicker priming Scene 7"}, {"4-5s": "Pedestrians sharpen as wireframe ghosts, AR visors glinting obliviously"}, {"5-6s": "Level to ground haze, Nova's trench silhouette materializes at frame's vanishing point, coat billow lingering into Scene 2 track"} ] }, "cinematography": { "lighting": "Desaturated neon primaries with volumetric god rays slicing haze for ethereal isolation, rain speculars adding dynamic highlights consistent across wet surfaces", "tone": "Oppressive immersion yielding to rebellious spark—global cyber-noir dread laced with glitch anticipation, flowing seamlessly to Nova's personal emergence" } }, "scene_2": { "global_cinematography": "Ultra-realistic Hollywood cyberpunk thriller in the vein of The Matrix (1999) and Blade Runner 2049 (2017), shot on Arri Alexa LF with anamorphic lenses for widescreen 2.39:1 aspect ratio, 24fps for fluid motion, desaturated palette dominated by cool blues, greens, and high-contrast neon reds piercing perpetual smog-choked night. Consistent VFX pipeline: Procedural green code cascades, photorealistic cybernetic augmentations with subsurface scattering, physics-based rain and particle simulations. Lighting paradigm: Volumetric god rays through haze, practical lens flares from holograms, rim lighting on metallic surfaces for depth. Sound integration: Pulsing industrial synth score with digital glitches, rain patter syncing to code interference, metallic echoes underscoring dialogue. Transitions: Seamless glitch wipes or matrix symbol dissolves ensuring narrative continuity, each scene's final beat priming the next for unbroken tension flow. Continuity directive: Scenes chain via lingering elements—rain droplets from prior shots persisting, Nova's silhouette echoing across cuts, HUD overlays threading flashbacks to present, escalating glitch distortions building to climax rupture—maintaining spatial and temporal cohesion in Neo-Tokyo's underbelly.", "shot": { "composition": "Low-angle tracking push with 50mm anamorphic prime on Arri Alexa LF, heroic distortion compressing background alley into claustrophobic funnel, foreground rain blur veiling initial fog for continuity from Scene 1 descent", "camera_motion": "Fluid forward Steadicam arc from lingering Scene 1 haze, subtle left profile tilt to frame Nova against graffiti wall, pulling back slightly to hold environmental depth into Scene 3 orbit" }, "subject": { "description": "Nova, 30s hybrid rebel with scarred synthetic pallor, cropped black hair rain-matted, holographic irises scanning with latent data flickers; sleek titanium limbs rune-etched in dormant blue", "wardrobe": "Sodden black trench coat with frayed hems from Scene 1 debris scatter, high collar shadowing jawline for motif continuity" }, "scene": { "location": "Graffiti-choked alley continuation from Scene 1 street convergence, Neo-Tokyo underbelly", "time_of_day": "Eternal neon-dusk syncing global palette", "environment": "Fog banks rolling from industrial vents as Scene 1 smog extension, wet cobblestones rippling with residual aerial rain patterns" }, "visual_details": { "action": "Nova materializes from Scene 1's terminal haze, striding assertively into sodium glow with metallic glint, coat hem dragging puddles to splash forward—teasing Scene 3 facial trace", "props": "Luminescent 'GLITCH THE SYSTEM' graffiti echoing from Scene 1 signs, overhead hover-traffic hum persisting from aerial hum", "action_sequence": [ {"0-1s": "Fog swirl from Scene 1 yields Nova's silhouette, boot first impacting puddle"}, {"1-2s": "Full stride forward, coat hem trails iridescent wake linking to blood drip in Scene 9"}, {"2-3s": "Titanium forearm catches neon, runes sequential-pulse awakening blue continuity"}, {"3-4s": "Holographic eyes iris-scan, reflecting alley code fragments priming Scene 4 overlay"}, {"4-5s": "Rain beads contour synthetic skin, parting at seams for Scene 3 macro journey"}, {"5-6s": "Profile lean against wall, vapor breath hangs, posture straightening into Scene 5 OTS"} ] }, "cinematography": { "lighting": "Harsh sodium sidelight rimming form per global motif, cool rune fill softening human remnants, prismatic rain refractions tying to Scene 1 aerial streaks", "tone": "Defiant grace in simulated decay—cyber-noir intimacy building personal stakes, camera arc ensuring spatial flow to close-up revelation" } }, etc etc etc up to scene 16. you got the point

NFK

3,351,561 views • 9 months ago

"Hah - generative ai can't even make an image of a hand with the right number of fingers.." "Stop pushing this slop" It's way past the point now where it must be clear to everyone, that generative ai is here to stay AND that the quality will continue to increase. I've been talking about this trajectory for years now, and I've been working towards finding ways to combine the strength of these models, with the best of what I love about "old-school" creation. Building with my hands, moving a pencil across the paper and seeing shapes emerge, moving a building slightly to the right to get just that composition I had in mind. Being fully immersed in a scene I'm building in VR, being inspired by the immersion to take the story in a new direction. For years I've been talking about how powerful the combination of 3d and generative ai is, be it traditional 3d, SDF volumes in Dreams or Gaussian splats - with experiments around using V2V as a "render pass" or with experiments around realtime ai. Enough talk you might think, where's the proof? It's all around us these days honestly and here's a small test I did during some OOO. Blender MPC + Fable - a pretty powerful combination! With a bit of Google Omni Fast on top as a "render" pass. What do you think of where this is heading? Hopeful, disheartened, inspired or the opposite? Can you imagine working with tools like this in a way where we still retain the human "spark" and the creative nerve that makes each persons creation unique?

Martin Nebelong

49,109 views • 20 days ago

The Social Network sequel also won’t have David Fincher. His directing style great combo with Sorkin’s dialogue-heavy script. Social Network’s opening scene is perfect example: it’s ~5 minutes and Fincher famously did 99 takes. It shows Zuck getting dumped in a Harvard bar (and with some creative licenses, how that leads to Facebook). Sorkin — who will direct the sequel — talked about how the scene came together: ➡️ "An average screenplay is about 120 pages long. My screenplays have higher page counts because there’s more dialogue and less action. By the rules of screenplay format, dialogue takes up more room on the page and less time on the screen than action (which takes up less room on the page and more time on the screen). "The Social Network" was 178 pages. And the studio said, “OK, the first thing you’ve got to do is figure out a way to cut 30 pages from this.” And David said, “I don’t think so.” […] He came over to my house with his iPhone set on stopwatch mode, and he said, “I want you to read the entire script out loud for me, at the pace you heard it in your head when you were writing it, and I’m going to write down the timing of each scene.” So that opening scene...with Jesse Eisenberg (Zuck) and Rooney Mara. I read it and it was 7 minutes and 22 seconds. In rehearsal, Jesse and Rooney would rehearse the scene, David would say great, and he would give them a couple of notes and always end with, “But this scene is 7 minutes and 22 seconds long, and you’re doing it at 7 minutes and 40 seconds. So I don’t care how, but you’re going to have to talk faster somewhere, because I promise you, this scene plays best at 7 minutes and 22 seconds.” ⬅️ The final cut was quite a bit shorter than 7 minutes and 22 seconds. But one thing stayed consistent throughout the shoot: Fincher told dozens of extras in the bar scene to keep the volume of their chatter high, as they would on any night out. This forced Eisenberg and Mara not only to speak faster. But also louder, increasing the scene’s intensity. IT WORKED REAL WELL!! *** Full interview with Sorkin here:

Trung Phan

57,943 views • 2 months ago