正在加载视频...

视频加载失败

AI motion capture just got scary good Kling 3.0 upgraded motion control that keeps your character identical across a 5-minute sequence -> same face, same walk, same everything Hollywood-grade stability finally here Here's how: credits phil.franco

36,284 次观看 • 5 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Made with kling 3.0 on Higgsfield AI 🧩 Video Prompt : Use the uploaded reference image as the exact identity reference for the subject. Create a hyper-realistic live IPL television broadcast crowd-shot sequence during a high-energy playoff cricket match in a packed Indian stadium. The subject from the reference image must remain fully consistent throughout: same face, same hairstyle, same outfit, same skin tone, same lighting, same seating position, and same crowd environment. Preserve identity accuracy strongly across the entire sequence. The video should feel exactly like a genuine Star Sports IPL crowd cutaway captured during a real live match — NOT cinematic, NOT influencer-style, NOT vlog-style. Show realistic live broadcast camera behavior: quick crowd cutaways, natural stadium lighting, authentic TV zoom lens movement, slight camera shake, realistic audience reactions, energetic IPL atmosphere, LED advertisement boards, match scoreboard overlays, cheering fans around the subject. The subject should react naturally to the match: smiling, clapping, looking tense during close moments, celebrating boundaries/wickets, and occasionally looking toward the field. Maintain realistic Indian stadium ambience with thousands of spectators, team jerseys, flags, chants, floodlights, and authentic IPL playoff energy. Ultra realistic skin texture, natural motion, realistic hair movement, accurate facial consistency, broadcast-quality detail, shallow depth of field, true live sports telecast aesthetic, 4K realism, highly detailed crowd environment. Prompt: cinematic lighting, music-video style, slow motion, overacting, beauty filter, influencer aesthetic, vlog framing, AI face distortion, cartoon look, unrealistic expressions, fantasy colors, excessive blur, duplicate faces, identity drift, studio lighting, posed acting, fake crowd.

Shahid Wani

139,457 次观看 • 2 个月前

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 次观看 • 11 个月前

THE DEPTH MAP TRICK THAT FIXED DANCE ACCURACY IN SEEDANCE 2.0 Feed the model a video of someone dancing and it tries to interpret everything- the person, the clothes, the lighting, the room, and somewhere in there, the movement. Feed it a depth map and there's nothing left to interpret but the motion. Most creators trying to transfer a dance to a character reference the source footage directly, then wonder why the choreography drifts. The problem isn't the model - it's that you handed it ten variables when you only wanted one. Here's the workflow 1. Lock the character reference in GPT Image 2 first -face, build, costume, so identity holds independently of whatever motion gets applied to it 2. Convert the source dance footage into a depth map instead of using the raw video -this strips out the original performer's appearance, clothing, and environment entirely 3. Feed the depth map as the motion reference and the character sheet as the identity reference- two separate inputs doing two separate jobs, not one input trying to do both 5. Let the depth map carry only spatial movement -the model receives body position and momentum with no competing information about who's moving or what they look like 6. Keep the character and motion inputs isolated throughout - the moment you mix appearance data into the motion reference, the model starts negotiating between two identities Why this works • Raw footage passes the model everything at once- performer, wardrobe, room, lighting -and the choreography competes with all of it for attention • A depth map is pure spatial information, so the only thing left to transfer is movement • Separating identity from motion means the character can stay locked while the dance stays accurate - normally you're trading one for the other • The accuracy gain isn't the model getting better, it's the model getting fewer decisions to make Use cases: ⁃ Dance and choreography transfer onto original characters ⁃ Motion capture-style workflows without motion capture ⁃ Any sequence where a specific movement needs to survive intact ⁃ Character showcase content built on existing performance footage The character sheet answers who's dancing. The depth map answers how - and keeping those two questions separate is the whole trick.

Nexlow

84,184 次观看 • 19 天前

(Yes i know the movement is exaggerated, its just to show the motion, it also doeant move like jello when in vr) -Left is before, Right is after- Ok wall of text time for the nerds Literally everything is the same on both sides except the position of 1 bone (made sure that even after adjusting, the collider would still be in the same spot) The reason for this change is because she wanted to be able to clap her ass in vr, and last time i did this all i did was add a toggle for a hidden bone that moved, and its angle would trigger the sound. Which was because that with my previous position, the cheeks could never really meet in the middle when in motion unless you just pushed them together with your hands, and i wanted to just have the sound actually be triggered by them colliding with each other This is also the first time ive used endbones to help drive the movement, so now thanks to that and all my previous experimenting with squishing as well, I can now get bones to move exactly how i want them to move when in natural motion. And i know a lot of people are gonna ask me to teach them how or explain my thought process but i literally just go off feeling lmao. Like, i just kinda visualize in my head how i want the motion to look, and then i just kinda know where to put everything, so im not even sure how to even start explaining rip. I have the things that artist want where they can just make the shit in their head exactly how they envisioned it lol I have a few more ideas i want to try, so we'll see if theyre good enough to get a tweet lol Thank you for coming to my ted talk

Pixel

27,240 次观看 • 1 年前

this effect is all over tiktok right now and nobody's explaining how to actually do it properly... the 3d balloon character thing. where someone turns into a shiny inflatable version of themselves that still moves and talks. looks pretty smooth in feeds. the workflow is stupid simple once you see it. step 1: take any photo. drop it into an image gen tool (nano banana pro). prompt it with something like "make the person in the photo a plastic blow up balloon character with a shiny surface. keep the face details as 3d balloon details including the person in the background. don't change background" that's it for the image. don't overcomplicate the prompt. shorter = more consistent results. (learned this after wasting like 2 hours trying to get "perfect" prompts that kept giving me garbage) step 2: take that balloon image + your original video and drop both into kling motion control. prompt: "turn the motion and detailed mouth movement of the video to the setting of the image" that's literally it. kling maps the motion from the real video onto the balloon character. mouth moves. head turns. expressions transfer. the whole thing renders in a few minutes. the result looks like a $500 custom animation and costs you maybe $0.30 in kling credits. people are getting 500k+ views with these because the scroll-stop factor is insane. nobody expects to see a shiny inflatable version of someone giving a real speech or doing a product review. the play here is obvious btw. run this for client content (mix with the hook and real body, check the results yourself) or use it on your own faceless channels as a hook pattern before the algo catches up...

KNOX

25,773 次观看 • 5 个月前