Loading video...

Video Failed to Load

Go Home

I am blown away 🤯. Check this out! CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models TL;DR: "To enable broader exploration of dynamic scenes, our model can generate new video clips of the same scene based on previously generated content and user-provided camera trajectories. This approach maintains...

12,633 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Wonderland: Navigating 3D Scenes from a Single Image Contributions: • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.

MrNeRF

52,849 views • 1 year ago

Depth Any Video with Scalable Synthetic Data AI physicists and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.

MrNeRF

27,428 views • 1 year ago

I’ve used all the recent GenAI video models extensively & here’s my 2¢: 🎬 Runway Gen3 Alpha - best image quality & motion for text-to-video & embedded words. Great at prompt travel changes over the course of 10 sec. And I’m super bullish on how gen3 will evolve, hopefully adopting the features listed below. Kling - best quality for image-to-video with prompt control, like eating food. Great clip extension that accounts for character (ie walking stride) & camera movement (speed & angle), rather than just using final frame. But it’s limited availability & Chinese native language is limiting. Used for Spider-Man video below (via Midjourney). LumaLabs - best for keyframe start & end control (it can not be overstated how important this is. other services should add it ASAP!) and their high dynamic action movements are really fun. Luma was used in my viral Multiverse of Memes video. PikaLabs - they haven’t gotten as much attention as others lately. But they did update their video model a few weeks ago and it looks great. Also, they are notable for their unique & AWESOME features, like video in-painting & out-painting. My perfect AI video platform would have the following features: 1) Gen3’s quality, prompt control & text embedding. 2) KLing’s image-to-video quality, prompt control & clip extension quality. 3) Luma’s multi-keyframe control & dynamic movement ability. 4) Pika’s inpainting & outpainting ability. And a video-to-video (aka next-gen Runway gen1) could be a game changer, too. It’s an exciting time to be alive 🫶 Who will get there first? 🔉🔉

Blaine Brown

26,535 views • 2 years ago

A cinematic AI-powered visual experience showcasing how imagination can be transformed into stunning digital worlds From futuristic environments and creative fashion visuals to seamless character transformations every scene blends technology creativity, and storytelling into one immersive piece. Made with MiniMaxH3 on WeryAI Wery Prompt: A stylish young South Asian woman with long dark hair, natural glamorous makeup and an elegant modern outfit. Keep her face, hairstyle, outfit and appearance consistent throughout the entire video. Scene 1 — 0–4s: She sits at a sleek modern workstation in a futuristic studio. She looks directly into the camera and says naturally: “What if you could turn your ideas into reality in just a few clicks?” Natural facial expressions and hand gestures. Slow cinematic camera push-in. Perfect lip-sync. Scene 2 — 4–7s: She turns to her computer and opens WeryAI. She types a creative idea and says: “That’s exactly what I love about WeryAI.” The interface responds with elegant AI processing animations and glowing digital effects. Scene 3 — 7–11s: Her simple idea transforms into stunning AI-generated visuals: a fashion scene, futuristic city, cinematic character, premium product advertisement and social-media video. She looks impressed and says: “Give it an idea, and watch it come to life.” Use smooth cinematic transitions and dynamic camera movement. Scene 4 — 11–15s: She turns back toward the camera with a confident smile and says: “WeryAI. Your imagination, powered by AI.” The camera slowly pulls back as the WeryAI brand reveal appears with a clean futuristic glow. Audio: Natural confident female voice, warm energetic delivery, subtle futuristic background music and soft cinematic sound effects. Keep dialogue clear and prominent. Lip-sync: The woman must visibly speak every line with accurate lip synchronization, natural mouth movement, blinking and facial expressions. Visual quality: Photorealistic, cinematic 4K, realistic skin, smooth camera movement, premium lighting, shallow depth of field, consistent character. No subtitles, no distorted face, no warped hands, no extra fingers, no flickering, no random text, no watermark.

Calira

12,635 views • 7 days ago

Two cinematic prompts for Seedance 2.0 right here 👇 [STYLE + CAMERA + ATMOSPHERE] Ultra-photorealistic cinematic 15-second action sequence from a grounded action thriller film. Nighttime at an abandoned industrial train yard. Old trains, rusted tracks, broken platforms and warehouses. Strong moonlight and industrial lighting. The visual language is gritty and realistic with hyper-realistic practical destruction. Camera is handheld-dynamic with urgent moves. [IMAGE REFERENCES] No reference images provided. Generate a fully consistent lead character: a male fugitive in his mid-30s, short hair, intense expression. He wears a dark jacket, jeans and boots. Maintain exact same appearance and clothing throughout the clip. [TIMELINE SECOND BY SECOND] 0-3s: [Dynamic handheld tracking shot from side-rear] The man runs across the train yard as a freight train derails violently after hitting a collapsed section of track. Massive train cars tip over and crash into the ground, sending metal, wood and cargo flying. The camera tracks him urgently as the destruction spreads. 3-7s: [Slow continuous orbiting camera moving around the frozen destruction] At the peak of the derailment, everything freezes in perfect realistic physics. Huge train cars, twisted metal, wooden debris and clouds of dust hang motionless in the air with accurate weight and trajectories. The camera performs a smooth, continuous orbit through the frozen chaos, moving between large suspended train cars and passing close to sharp metal fragments while the man remains visible. 7-15s: [Dynamic tracking shot as time resumes] At the 7-second mark time snaps back to forward motion with a slight realistic temporal residue. Some outer debris shows a very subtle reverse lag before continuing. The man has used the frozen moment to dive behind a concrete barrier. When motion resumes, several large pieces of metal miss him. He stays low and runs toward an open warehouse as debris finally crashes down. The camera tracks with him through the dust. [STYLE & QUALITY BOOSTERS] Photorealistic 8K, hyper-realistic rigid body destruction with correct material properties (metal, wood, train parts), perfect frozen mid-air physics, seamless transition from freeze to resume, natural dust interaction, heavy cinematic motion blur on fast debris, stable character performance, movie-level practical destruction VFX quality.

TechHalla

16,316 views • 1 month ago