Загрузка видео...

Не удалось загрузить видео

На главную

This is character consistency with Ray3. Subjects maintain identity as they travel through environments or uphold features within spatial changes. Characters remain clear and coherent across every frame, preserving fidelity so scenes retain realism, continuity and visual depth.

22,697 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Depth Any Video with Scalable Synthetic Data AI physicists and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.

MrNeRF

27,428 просмотров • 1 год назад

📖THE STEP MOST CREATORS SKIP IS WHY THEIR AI ANIMATION LOOKS INCONSISTENT Consistency across clips doesn't come from prompting — it comes from the reference image. The pipeline, step by step: ▪ Start with ChatGPT Image 2 — generate a full character design sheet first, not just a single frame. Multiple angles, expressions, and outfit variations in one image keeps the character consistent across every scene ▪ Build a storyboard inside ChatGPT Image 2 as well — define each shot, camera angle, action, and mood before touching Seedance at all. This is the step most people skip and it's the reason clips look disconnected ▪ Define a color palette and lighting mood early — golden afternoon light, soft warm tones, dramatic shadows. Lock those values and repeat them across every prompt ▪ Take each storyboard frame into Seedance 2.0 as the reference image — one frame becomes one clip ▪ Write the Seedance prompt around the character action, not the scene description. The scene is already in the image. The prompt handles motion, camera behavior, and timing ▪ Keep clip duration between 4-6 seconds per shot — shorter clips give more control over pacing and reduce motion drift on character faces ▪ Match camera movement type across consecutive clips — if one shot dollies in, the next should hold or pull back, not dolly again The consistency across these frames comes from the character design sheet, not from luck. Seedance reads the reference image and the prompt together — if the reference is detailed enough, the output stays on-model. This video was created by ALOKXMEHTA 📥 tomorrow: the exact ChatGPT Image 2 prompt structure used to generate a multi-angle character design sheet like this one 🔖One article covers the entire workflow — it is pinned below, do not scroll past it.

Zentrix⌚️

14,015 просмотров • 2 месяцев назад

Would you underestimate her just because she wears a school uniform? GPT Image 2 + Seedance 2.0 on Sjolt Try Canvas: prompt Character Identity Lock (Highest Priority): Use the exact same young East Asian woman from the provided reference character sheet. Preserve 100% identical facial features, face shape, eye shape, nose, lips, skin tone, hairstyle, hair color, proportions, and overall identity throughout the entire video. Do not redesign, reinterpret, or substitute the character. She must remain instantly recognizable as the same person from the reference image. She has shoulder-length wavy silver-gray hair with subtle blue undertones, bright expressive eyes, fair skin, and a confident slight smile that naturally transitions into a focused, determined combat expression. She wears the identical navy blue Korean high school uniform blazer over a gray sweater vest, white collared shirt, striped tie, and matching school skirt from the reference character sheet. Video Prompt: A cinematic, hyper-realistic action sequence inside a chaotic South Korean high school classroom. The classroom is filled with overturned desks, scattered chairs, flying notebooks, broken pencils, and papers drifting through the air. Bright natural daylight streams through large classroom windows, creating realistic highlights, soft shadows, and cinematic contrast. The young female student moves with incredible speed, confidence, and precision as she expertly defends herself against multiple aggressive male students wearing matching Korean school uniforms. Every movement is fluid, athletic, and grounded in realistic martial arts choreography. The camera remains highly dynamic, featuring cinematic handheld tracking shots, fast push-ins, orbit shots, dramatic slow-motion moments, whip pans, low-angle hero shots, and close-up impact shots. Capture rapid combinations of punches, clean high kicks, evasive footwork, parries, elbow strikes, blocks, and throws. Desks slide across the floor, chairs topple over, and dust particles catch the sunlight, emphasizing the intensity of the action. Maintain a high shutter-speed action-photography aesthetic with crisp motion detail, subtle motion blur only during extremely fast movements, physically accurate body mechanics, realistic cloth simulation, natural hair physics, authentic facial expressions, and believable impact reactions. Keep the camera frequently returning to sharp close-ups of her face to reinforce character continuity and emotional intensity. Her silver-gray hair flows naturally with every movement while her determined eyes remain locked on her opponents. Photorealistic cinematic quality, 4K HDR, ultra-detailed skin textures, realistic lighting, volumetric daylight, physically based rendering, shallow depth of field during close-ups, blockbuster Korean action film aesthetic, empowering heroine energy, consistent facial identity throughout every frame, no face drift, no character variation, no animation-style exaggeration.

Sharon Riley

26,184 просмотров • 2 месяцев назад

GPT-Image-2 + Seedance 2.0目前已成AI视频标配 甚至可以根据给定图片推导过去和未来,制作storyboard,然后生成视频 使用方法: 1️⃣ 随便找一张图 2️⃣ 给以下提示词,然后制作storyboard 用以下提示词👇: Create a 3×3 cinematic storyboard grid based on the uploaded reference image. Use the uploaded image as the central moment of the story: Frame 5 must represent the exact “t” moment, matching the subject, scene, mood, composition, costume, environment, lighting style, and emotional tone of the reference image. The storyboard must show what happened before and after this moment as a time-based visual timeline. FRAME STRUCTURE: Frame 1: t-30: Establishing shot, the wider environment before the main event begins. Frame 2: t-10: The subject approaches or prepares for the key moment. Frame 3: t-5: Tension builds, body language and atmosphere lead toward the reference image. Frame 4: t-1: Final instant before the reference image, close emotional or action transition. Frame 5: t: Recreate the uploaded reference image as the central key frame. Frame 6: t+1: Immediate reaction or continuation right after the key moment. Frame 7: t+5: Alternate angle showing the consequence of the moment. Frame 8: t+15: Candid transition frame, natural movement, emotional aftermath. Frame 9: t+30: Strong final cinematic frame that clearly resolves the scene. STYLE: Ultra-realistic cinematic storyboard, 3×3 grid layout, cohesive visual tone across all frames, consistent character identity, consistent costume, consistent environment, cinematic lighting, shallow depth of field, realistic camera angles, natural motion continuity, no text labels, no numbers, no arrows, no captions inside the image. 3️⃣ seedance2.0 一键成片

Jason Zhu

34,160 просмотров • 4 месяцев назад

Wonderland: Navigating 3D Scenes from a Single Image Contributions: • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.

MrNeRF

52,849 просмотров • 1 год назад

EVERYONE PROMPTS THE ACTION. ALMOST NOBODY LOCKS THE IDENTITY — WHICH IS WHY TWO-CHARACTER SCENES FALL APART. Two freerunners racing across Tokyo rooftops, eight cuts, corkscrews over a rooftop gap at the end. The parkour is the easy part. Keeping them two separate people who never blend into each other is the part that actually breaks. Here's the full prompt built that way. Attach two reference photos as image_1 and image_2, and the same structure works for any multi-character action piece: FORMAT: 15 seconds, 16:9, 1080p, 8-cut cinematic ultra-advanced parkour footage. CHARACTERS: Two realistic individuals from image_1 and image_2. Use the attached images as absolute character references, and fully maintain the facial features, hairstyles, hair colors, skin textures, body types, height differences, outfits, color schemes, and age appearances of each person across all cuts. No altering into different people, face swaps, outfit changes, hairstyle changes, or mixing of the two individuals' features. SETTING: A sunny modern Japanese city reminiscent of Tokyo, Shibuya, and Yokohama — rooftops, alleys, staircases, railings, pipes, concrete walls. The two protagonists, as equals, race through at high speed running side by side, following, crossing paths, and coordinating. CUTS: 1. (00:00–00:01.60) Low-angle rear tracking. Both accelerate side by side and simultaneously kong vault over separate obstacles. 2. (00:01.60–00:03.40) Front low-angle. One wall runs the left wall, the other the right, then tic-tac to cross in midair and land on opposite rooftops. 3. (00:03.40–00:05.20) Lateral tracking. Consecutive precision jumps, then cat leaps to grab and climb a high wall. 4. (00:05.20–00:07.20) Rooftop tracking. The leader dash vaults, the trailer websters over the gap, then they swap front and back positions. 5. (00:07.20–00:09.20) Overhead moving camera. Both dive roll, then run side by side to speed vault a long railing. 6. (00:09.20–00:11.30) Handheld retreating from the front. One underbars, the other side flips, conquering the obstacle simultaneously. 7. (00:11.30–00:13.20) Drone from diagonal rear above. Both palm spin off left and right walls, kong vault, accelerate into the final jump. 8. (00:13.20–00:15.00) Climax. Both leap a large rooftop gap, each doing a corkscrew, camera circling them in midair as they land on separate rooftop edges — then run side by side into the distance. QUALITY: Live-action film quality. World-championship-level smooth freerunning. Realistic center-of-gravity shifts, muscle movement, natural landing impacts, swaying hair and clothing. Sharp background, natural motion blur only during high-speed movement. PROHIBITED: Facial distortion, altering into different people, face or body swaps, outfit changes, hairstyle changes, body type changes, limb multiplication, duplicates, body fusion, penetration, warping, floating, unnatural landings, anime style, CG style. A few things worth noticing about why it's built this way: The character block does identity work three separate times — the reference images, the "fully maintain" list, and the prohibited list at the end. That redundancy isn't padding; each one closes a different door the model tends to walk through. The prohibited list names the exact failure modes — face swaps, body fusion, limb multiplication. Telling the model what not to do is more effective here than describing what you want, because these are the specific ways two-character scenes collapse. Every cut assigns each person a distinct action — one wall runs left, the other right; one underbars, the other side flips. Giving them separate roles keeps them functionally two people, so the model can't average them into one. And the cuts are individually timed and framed. Long continuous motion is where identity drift creeps in — breaking it into eight discrete shots gives the model less room to blend them. Made in Seedance 2.0.

Nexlow

114,432 просмотров • 2 месяцев назад

Sora 2 + n8n is absolutely insane 🤯 This n8n automation generates entire UGC campaigns with the same AI creator across unlimited videos. All from one Airtable form. Perfect for DTC brands & agencies who need brand consistency in their AI ads without hiring real creators. Why this matters: Every AI video tool gives you a random person each time. You can't build multi-video campaigns because your "creator" changes in every clip. Sora 2 consistent characters solves this: Same AI creator → Different scenes → Unlimited videos The n8n workflow: → Fill out Airtable form once (select your character, describe scenes, choose quantity) → Claude AI generates professional Sora 2 prompts automatically → Sora 2 renders videos with your consistent character → Videos auto-upload to ImageKit CDN → Everything tracked in Airtable with shareable URLs No manual prompting. No file management. No different people in every video. What you can create: → 3-part testimonial series with the same person → Before/during/after transformation campaigns → Product tutorial sequences that feel cohesive → Entire ad creative libraries with your "brand ambassador" Track everything in Airtable: → Video status (queued → generating → complete) → Shareable URLs for each clip → Scene descriptions and prompts → Production-ready in 5-10 minutes Built 100% in n8n + Airtable. Want the complete template? > Comment "SORA" > Like this post And I'll send it over (must be following so I can DM)

Mike Futia

19,054 просмотров • 10 месяцев назад

A 24-year-old built two AI girls with Claude and now clears $21,800 a month from them. The build took 15 days. He trained separate LoRAs for both girls, locked their identity seeds, and kept small imperfections on purpose: a loose strand of hair, tiny skin marks, slightly uneven framing. Perfect symmetry gets flagged. Small inconsistencies make them look real. He posts 5 times a day across TikTok, Instagram and X. Morning routines, gym sessions, mirror videos, outfit changes, pool clips, and videos of the two girls together. The content is designed so they look like two real friends who actually live in the same world. The smartest part is that the accounts interact with each other. One girl comments on the other’s posts, appears in her videos, and references things they supposedly did together. Followers stop seeing them as two AI models and start following the relationship between the characters. Within three months they crossed 312,000 followers combined and started receiving hundreds of DMs every night. The private channel sits at $25 a month, while an AI memory agent keeps track of every conversation, previous message, favorite post, and personal detail each follower has shared. Replies come back in under 30 seconds. The agent checks the user's previous conversations before answering, so the response feels consistent with the personality of the girl they are talking to instead of sounding like another generic AI chatbot. By the end of month three, the two accounts were generating $13,900 from subscriptions and private chats, another $5,700 from brand deals, and around $2,200 from digital products. The brands came after the audience started growing: clothing companies, beauty products, fitness brands, and lifestyle products wanted access to the same audience that was already following the two characters every day. The Claude stack that locked them: 1Full identity, personality, lighting, camera style and body proportions locked into separate character systems. 2Separate LoRAs trained only on each girl's approved character frames. 3Apartment, bedroom, gym and outdoor locations generated once and reused to keep the world consistent. 4Every video built around natural movement, imperfect framing and small variations instead of polished AI-perfect shots. 5Memory agent connected to the conversations so both girls remember what followers previously said. 6Upscaling, face consistency and final post-processing before everything goes live. The first girl brings people into the account. The second gives them another character to follow, another story to watch, and another reason to come back. The content gets them interested. The relationship between the two characters keeps them watching. The memory agent turns that attention into recurring revenue.

genuenci

678,031 просмотров • 1 месяц назад