🔥 AI video creation is finally evolving beyond basic... motion prompts. Try out Depth Map Control right here: Historically, the hardest part of the process has been maintaining scene consistency: • camera movement • object placement • spatial relationships Thanks to PixVerse Depth Map Control, you can now grab depth data from a reference video and merge it with a fresh image to explore highly directed video generation. The final outcome? Much smoother camera tracking. Highly cohesive visual scenes. It is a powerful workflow that hands creators actual authority over how their AI visuals shift and flow over time. 🎬show more

Elara Quinn
12,160 views • 2 days ago
I handed a dance video over to AI... Try... this exact workflow yourself: I used PixVerse Depth Map Control to seamlessly map the original choreography onto a completely fresh character. The visual outcome is incredibly fascinating 👀 Instead of generating a scene from absolute zero, this specific setup helps you lock down crucial details from your base clip: • movement • body position • spatial structure Your source footage dictates the precise action, while a single reference picture establishes the new aesthetic. It is a brilliant way to level up your AI video experiments — granting you far more precision over how your animations move and shift.show more

LX™
59,128 views • 3 days ago
OpenArt just launched Seedance 2.0 for Teams and Enterprise... It delivers the kind of camera control, reference depth, and cinematic storytelling most AI video tools still struggle to offer You can use up to 9 image references, 3 videos, and 3 audio files to build multi-shot scenes with far more control than a normal prompt 10 prompts worth saving:show more

Amira Zairi
40,069 views • 4 months ago
THE DEPTH MAP TRICK THAT FIXED DANCE ACCURACY IN... SEEDANCE 2.0 Feed the model a video of someone dancing and it tries to interpret everything- the person, the clothes, the lighting, the room, and somewhere in there, the movement. Feed it a depth map and there's nothing left to interpret but the motion. Most creators trying to transfer a dance to a character reference the source footage directly, then wonder why the choreography drifts. The problem isn't the model - it's that you handed it ten variables when you only wanted one. Here's the workflow 1. Lock the character reference in GPT Image 2 first -face, build, costume, so identity holds independently of whatever motion gets applied to it 2. Convert the source dance footage into a depth map instead of using the raw video -this strips out the original performer's appearance, clothing, and environment entirely 3. Feed the depth map as the motion reference and the character sheet as the identity reference- two separate inputs doing two separate jobs, not one input trying to do both 5. Let the depth map carry only spatial movement -the model receives body position and momentum with no competing information about who's moving or what they look like 6. Keep the character and motion inputs isolated throughout - the moment you mix appearance data into the motion reference, the model starts negotiating between two identities Why this works • Raw footage passes the model everything at once- performer, wardrobe, room, lighting -and the choreography competes with all of it for attention • A depth map is pure spatial information, so the only thing left to transfer is movement • Separating identity from motion means the character can stay locked while the dance stays accurate - normally you're trading one for the other • The accuracy gain isn't the model getting better, it's the model getting fewer decisions to make Use cases: ⁃ Dance and choreography transfer onto original characters ⁃ Motion capture-style workflows without motion capture ⁃ Any sequence where a specific movement needs to survive intact ⁃ Character showcase content built on existing performance footage The character sheet answers who's dancing. The depth map answers how - and keeping those two questions separate is the whole trick.show more

Nexlow
84,184 views • 17 days ago
Seedance 2.0 is now LIVE on the Pollo AI... App (iOS & Android) I’ve actually tried Seedance 2.0 on Pollo AI, and I can say it’s a solid upgrade. The motion feels more natural, and having audio and visuals generated together makes the whole output more cohesive. What stood out for me most is the level of control — being able to adjust performance, lighting, and camera really changes how you approach storytelling. 🔥 Plus, grab 60% OFF for a limited time on Pollo AI! Perfect for anyone diving into AI video creation.show more

Leonardo
31,641 views • 4 months ago
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 views • 1 year ago
Kling AI 3.0 is here And it’s a serious... step forward for AI video. This update isn’t about small tweaks. It’s about refinement, realism, and making AI video feel ready for real-world use. What stands out with Kling 3.0? Stronger visual quality. Movements feel smoother. Lighting looks more natural. Scenes feel intentional instead of generated. Better prompt understanding. You can describe complex scenes, moods, or camera directions, and Kling 3.0 interprets them with much more accuracy. Less trial and error. More usable results. Improved motion and consistency. Characters stay consistent. Objects behave logically. Shots flow more naturally from start to finish. Greater creative control. From product-style visuals to cinematic storytelling, Kling 3.0 gives creators the flexibility to move beyond simple clips and into structured, high-quality sequences. The difference is subtle until you see it, and then it’s obvious. AI video is evolving quickly, and Kling AI 3.0 shows how far the technology has come. It’s not just about generating video anymore. It’s about generating video that’s usable, polished, and ready for campaigns, storytelling, and branded content. If you’re building with AI, this is one of the tools worth paying attention to, best AI video model for pro & commercial productionshow more

Future Stacked
183,095 views • 5 months ago
Seedance 2.0 is insane... AI filmmaking is no longer... locked behind complex workflows, expensive tools, or regional limits. SJinn Agent now supports both Seedance 2.0 Pro and Seedance 2.0 Fast, giving creators a faster way to generate cinematic videos with more control over the final result. You can add image, video, and audio references to guide the direction, motion, style, and feeling of your videos, making the process more like directing And the best part: they’re offering 40% off, so this is probably the easiest time to test what high-level AI video creation can actually look like Prompt in first comment:show more

Amira Zairi
55,860 views • 2 months ago
1/ We've all been aware of the hype surrounding... Dreamina Seedance 2.0, and I finally got early access to the tool. And wow, I'm blown away. This is by far the best. It is starting to make AI video feel less like "generate a clip" and more like "direct a scene." What stood out to me is the level of control: camera motion, pacing, visual consistency, and the ability to build from multiple references inside one workflow. Some of the prompt directions that feel especially strong: - a busy modern city square during daytime. Suddenly, time freezes completely - a single continuous camera movement through a natural landscape that transitions through all four seasons in one shot - an underwater bioluminescent city waking up at dawn The big shift is this: One Prompt, Viral Remade. Edit Videos as Easy as Editing Photos. Dreamina Seedance 2.0 feels like a real step toward AI-native directing rather than just AI generation. Here are some examples 🧵:show more

Chubby♨️
75,394 views • 4 months ago
Gemini Omni's motion control is f*cking cracked i just... figured out how to turn 1 reference video into 50+ AI videos with the exact same movements... you have a video of someone eating, dancing, using a product, doing whatever complex motion you need. you feed it to Gemini Omni and it recreates that exact motion with a completely new AI character in literally one prompt i've tested this against Kling motion control and it's not even close. Kling falls apart the moment you try anything complex. eating scenes look weird, hand movements get mangled, anything multi-step breaks down completely. Gemini Omni handles all of it if you're still using kling motion control or paying creators to split test your videos, this replaces that entire workflow here's the thing though. you can't just prompt this out of the box. if you try to do motion transfer with default prompting you're going to get errors or the motion won't transfer properly. there's a specific prompting method that makes it work every time so i packaged up the whole system.. here's what you're getting: > full step by step video breakdown > how to find the best reference videos to use > the exact prompting system that allows for motion control transfer so you never get errors > the workflow for batching this out at scale (1 video → 50+) RT + reply "MOTION" and i'll send it over (must follow so i can dm)show more

Miko
58,182 views • 21 days ago
Kling 2.6 Motion Control is absolutely insane 🤯 Take... any reference video and transfer the exact motion onto an AI character: full-body sync, facial expressions, hand gestures, everything. All with just a few clicks. Perfect for e-comm brands and agencies creating AI video ads that don't look like AI. Here's the problem: AI-generated video ads still look robotic. The movements are stiff, the expressions are flat. Your audience clocks it as AI instantly and keeps scrolling. Kling 2.6 Motion Control fixes it: → Start with any reference clip (stock footage, existing UGC, motion reference) → Upload to Kling → Map the exact movement onto any AI character → Full-body motion, hand gestures, facial expressions—all transferred → Generate up to 30 seconds of video No stiff AI movements, no uncanny valley, no instant "skip this ad" reaction. What this unlocks: - Use one winning UGC motion → swap in different AI creators - Pull reference clips from anywhere → generate branded variations - Create dynamic AI video ads with real human movement - Test multiple "creators" without filming anyone new I recorded a quick walkthrough showing how to do this step-by-step. Want access? > Comment "KLING" > Like this post And I'll send it over (must be following so I can DM)show more

Mike Futia
25,179 views • 6 months ago
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,849 views • 1 year ago
Dreamina Seedance 2.0 is Officially here! Dreamina Seedance 2.0,... ByteDance’s AI-powered creative platform, lets creators easily transform ideas into high-quality videos. You can now edit videos like images, using up to 4 reference modalities (video, image, audio, and text) with precise control over visual effects, camera movements, and more. Why is this revolutionary? - 🔥 One-Prompt Video Editing: Edit videos seamlessly as easily as editing images. - 🎬 Multimodal Creativity: Combine images, video, audio, and text—up to 12 files at once. - 🎥 Remake Viral Content: Create high quality, professional-level videos with intelligent AI. - 🔧 Full Creative Control: Keep consistency in shots, typography, camera flow, and more. This isn’t just another video editing tool, it’s a one-stop AI workspace for all your creative needs. Ready to level up your content? Explore Dreamina Seedance 2.0 and start creating today. 🔗 #dreamina #seedance2 #seedream5 #dreaminatutorial #ai #aitools #aidesign #ecommercedesign #digitalmarketing #startupbusinessshow more

GitHub Projects Community
18,440 views • 4 months ago
Kling 2.6 Motion Capture is so good. It's a... huge leap forward. We've been able to do motion capture with AI for a while, but the quality bump finally made this approach useful. Here are the steps: - Get a reference video - record it or get a stock video - Grab the start frame of that video - Edit it using Nano Banana Pro. Ask to replace character and background. - Select the Kling 2.6 Motion Capture model in the video generator on Freepik (now Magnific) - Upload reference video + set the edited start frame - Video prompt can be simple e.g. "Gandalf dancing"show more

Martin LeBlanc
35,226 views • 7 months ago
The next leap in AI video isn't just better... visuals. It's having far more control over how those videos are created. Dreamina Seedance 2.5 is coming soon to the Dreamina platform, and what stands out isn't only the quality it's the workflow built around creators. Here's what's coming: • Multimodal input — combine up to 50 reference assets in a single generation. • Longer video generation — create continuous videos of up to 30 seconds. • Structured control — use white-model and green-screen references for more predictable results. • Multilingual creation — build content for international audiences with localization support. • Targeted refinement — edit specific parts of a video instead of regenerating the entire scene. These upgrades make Dreamina Seedance 2.5 feel less like another prompt-to-video model and more like a complete, controllable AI video production pipeline. Learn more: #dreamina #dreaminapartner #seedance #dreaminaseedance25show more

Md Riyazuddin
33,164 views • 6 days ago
This week is already so hot. 🔥 Massive release... from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."show more

Rohan Paul
12,761 views • 6 months ago
NVIDIA finally released Neuralangelo's source code! The model can... turn videos from any device into detailed 3D structures, fully replicating buildings, sculptures, or other real aworld objects or spaces virtually. Here's how it works: A model utilizes a 2D video with multiple angles of an object or scene. I selects frames from different viewpoints to understand depth, size, and shape. The AI creates an initial 3D representation, similar to a sculptor shaping a subject. The render is optimized to enhance details, like a sculptor refining texture. The outcome is a 3D object or scene suitable for virtual reality, digital twins, or robotics.show more

Lior Alexander
478,046 views • 3 years ago
I made a second version of the transformation video.... This time I moved the character into a real environment instead of using a plain white background. I also changed the opening and showed the final transformation result first. The overall flow feels much stronger now. Full workflow + prompts: The interesting part is that this type of video is much easier to make than it looks. You only need a base image, then generate the next keyframe through simple edits or outfit changes on the canvas. The entire transformation sequence is built from those keyframes.show more

underwood
313,847 views • 1 month ago
🔥 VIDU Multi-Entity Consistency Give Vidu 2/3 images and... it’ll turn them into a video—it’s pure magic! ✨ Your own characters interacting with objects and in the exact environment you want! Ads, movies… endless possibilities, and this is just the beginning! Thanks @Viduforhuman The future is a carrot! 🥕 Plus, how about grabbing any frame from a Vidu-generated video "from scratch" and throwing it into another AI video or image tool to push your project even further? For now, check out the comment below: I scaled up a frame with Magnific.ai and fed it into Runway to create a dynamic shot using full camera control. But fingers crossed I can soon use #ReCapture by Bisho & team to generate new shots from the same video!show more

Hungry Donkey 🥕
37,561 views • 1 year ago
Some updates on the multiview vistadream pipeline with Rerun!... Rerun came in extremely useful here, as being able to visualize depths at each stage of the pipeline allowed me to debug some nasty bugs. Since the last time, I was only working with a single image input. I've added in VGGT as my multiview pose + depth estimator. It works REALLY well for getting camera poses, but the depths are not that great. To try and fix that, I estimated depth maps from MoGeV2 for each of the views, and scale+shift aligned them so that they would match up to the confident sections of VGGT's depth predictions. You can see in the video just how much sharper the visualized 2d depth maps are! The biggest issue continues to be the multiview consistency 🫠 That's up next, along with actually training the Gaussian splat. Lots of work went into actually understanding inputs+outputs for VGGT. I had some funky bugs where the confidence values would all collapse to true I'm also really excited for this pipeline to use Difix3D+ Nvidia instead of Flux Inpainting, it seems like a better suited for a multiview pipeline.show more

Pablo Vela
29,904 views • 11 months ago
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,911 views • 3 months ago