Introducing “FlowCam: Training Generalizable 3D Radiance Fields w/o Camera... Poses via Pixel-Aligned Scene Flow”! We train a generalizable 3D scene representation self-supervised on datasets of raw videos, without any pre-computed camera poses or SFM! 1/nshow more

Vincent Sitzmann
88,503 görüntüleme • 3 yıl önce
F3D-Gaus: Feed-forward 3D-aware Generation on ImageNet with Cycle-Consistent Gaussian... Splatting Contributions: • We pioneer 3D-aware generation using generalizable feed-forward Gaussian Splatting representation, achieving significant efficiency and favorable rendering quality on monocular datasets. • We significantly advance the capability of pixel-aligned Gaussian Splatting representations by designing a self-supervised cycle training strategy specifically tailored for monocular datasets. • We further mitigate the artifacts of 3D-aware representations caused by large viewpoint shifts by introducing geometry-aware video priors.show more

MrNeRF
14,229 görüntüleme • 1 yıl önce
3D Gaussian Splatting for Real-Time Radiance Field Rendering paper... page: Radiance Field methods have recently revolutionized novel-view synthesis of scenes captured with multiple photos or videos. However, achieving high visual quality still requires neural networks that are costly to train and render, while recent faster methods inevitably trade off speed for quality. For unbounded and complete scenes (rather than isolated objects) and 1080p resolution rendering, no current method can achieve real-time display rates. We introduce three key elements that allow us to achieve state-of-the-art visual quality while maintaining competitive training times and importantly allow high-quality real-time (>= 30 fps) novel-view synthesis at 1080p resolution. First, starting from sparse points produced during camera calibration, we represent the scene with 3D Gaussians that preserve desirable properties of continuous volumetric radiance fields for scene optimization while avoiding unnecessary computation in empty space; Second, we perform interleaved optimization/density control of the 3D Gaussians, notably optimizing anisotropic covariance to achieve an accurate representation of the scene; Third, we develop a fast visibility-aware rendering algorithm that supports anisotropic splatting and both accelerates training and allows realtime rendering. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets.show more

AK
633,674 görüntüleme • 3 yıl önce
I am blown away 🤯. Check this out! CameraCtrl... II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models TL;DR: "To enable broader exploration of dynamic scenes, our model can generate new video clips of the same scene based on previously generated content and user-provided camera trajectories. This approach maintains dynamic capabilities, accurate camera control, and scene consistency throughout the extended exploration." "Our model enables precise camera control across diverse scenarios while preserving dynamic scene elements, e.g." "Our method can generate videos with strong 3D consistency, which enables high-quality 3D reconstruction using the camera-controlled videos." Contributions: 1) A systematic data curation pipeline for constructing a dynamic video dataset with camera trajectory annotations; 2) A lightweight camera control injection module and corresponding training strategy that preserves dynamic video generation capabilities while adding camera control effect; 3) A clip-wise autoregressive generation recipe that enables extended range exploration of generated scenes.show more

MrNeRF
12,633 görüntüleme • 1 yıl önce
Meta releases VGGSfM Visual Geometry Grounded Deep Structure From... Motion Structure-from-motion (SfM) is a long-standing problem in the computer vision community, which aims to reconstruct the camera poses and 3D structure of a scene from a set of unconstrained 2D images. Classical frameworks solve this problem in an incremental manner by detecting and matching keypoints, registering images, triangulating 3D points, and conducting bundle adjustment. Recent research efforts have predominantly revolved around harnessing the power of deep learning techniques to enhance specific elements (e.g., keypoint matching), but are still based on the original, non-differentiable pipeline. Instead, we propose a new deep SfM pipeline VGGSfM, where each component is fully differentiable and thus can be trained in an end-to-end manner. To this end, we introduce new mechanisms and simplifications. First, we build on recent advances in deep 2D point tracking to extract reliable pixel-accurate tracks, which eliminates the need for chaining pairwise matches. Furthermore, we recover all cameras simultaneously based on the image and track features instead of gradually registering cameras. Finally, we optimise the cameras and triangulate 3D points via a differentiable bundle adjustment layer. We attain state-of-the-art performance on three popular datasets, CO3D, IMC Phototourism, and ETH3D.show more

AK
96,527 görüntüleme • 2 yıl önce
📢Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single... Image📢 We directly regress neural parametric head models (NPHMs) from a single image — fast, stable, and significantly more expressive than classical 3DMMs such as FLAME. Face tracking & 3D reconstruction are often limited by the representational capacity of PCA-based face models. By lifting NPHMs to a first-class reconstruction primitive, we enable more accurate geometry, richer expressions, and finer animation control. Pix2NPHM obtains fast and reliable NPHM reconstructions on real-world data. Inference-time optimization against surface normals and canonical point maps can further increase fidelity. Key to successful and generalized training of our ViT-based network are: (1) large-scale registration of existing 3D head datasets, and (2) self-supervised training on vast in-the-wild 2D video datasets using pseudo ground-truth surface normals. Finally, we show that geometry-aware pretraining on pixel-aligned reconstruction tasks significantly outperforms generic visual pretraining (e.g., DINO-style features) in terms of generalization. 🌍 🎥 Great work by Simon Giebenhain, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Zhe Chenshow more

Matthias Niessner
37,978 görüntüleme • 8 ay önce
NeuRBF: A Neural Fields Representation with Adaptive Radial Basis... Functions paper page: present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The spatial positions of their neural features are fixed on grid nodes and cannot well adapt to target signals. Our method instead builds upon general radial bases with flexible kernel position and shape, which have higher spatial adaptivity and can more closely fit target signals. To further improve the channel-wise capacity of radial basis functions, we propose to compose them with multi-frequency sinusoid functions. This technique extends a radial basis to multiple Fourier radial bases of different frequency bands without requiring extra parameters, facilitating the representation of details. Moreover, by marrying adaptive radial bases with grid-based ones, our hybrid combination inherits both adaptivity and interpolation smoothness. We carefully designed weighting schemes to let radial bases adapt to different types of signals effectively. Our experiments on 2D image and 3D signed distance field representation demonstrate the higher accuracy and compactness of our method than prior arts. When applied to neural radiance field reconstruction, our method achieves state-of-the-art rendering quality, with small model size and comparable training speed.show more

AK
194,469 görüntüleme • 2 yıl önce
I've seen a lot of animatics lately that are... super detailed, essentially viewport previews of the final shot. But an animatic at it's core doesn't need to be anything fancy. Its main purpose is to plan and test the timing, camera angles, movement, and overall composition. With animatics you have to keep in mind the following: Pre-visualization: It allows to see a basic version of the animation or scene before committing to detailed work. Timing and pacing: An animatic helps identify how long each scene or shot should last, ensuring that the timing feels right. Planning: It helps with layout, camera angles, and transitions. By visualizing the shots, the team can ensure that the framing and overall design of the scenes work well in 3D space. Efficiency: It allows the team to test and fix any potential issues early on, like awkward movements or awkward pacing, before spending time on high-quality rendering or complex animation. In short, an animatic helps in conceptualizing the final animation by giving a low-res, rough version of the scenes, which guides the entire production process in terms of visual design, timing, and storytelling.show more

Voxyde
22,244 görüntüleme • 1 yıl önce
🚀 Introducing EgoExo Forge - built on top of... Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tunedshow more

Pablo Vela
35,919 görüntüleme • 1 yıl önce
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,992 görüntüleme • 5 ay önce
POV: you’re on a superhero movie set and suddenly... realize the entire building is lying flat on the ground 😂🎬 Made possible with 365 days of unlimited MiniMax H3 on Pollo AI, without having to worry about extra credits. Prompt: Create a highly realistic viral behind-the-scenes superhero filmmaking video inspired by the kind of practical forced-perspective stunt videos that go viral on TikTok, Instagram, and X — but make the concept, staging, characters, and visual execution original. The superhero is an original Superman-inspired flying superhero: a handsome adult male superhero wearing a premium blue-and-red superhero suit with a flowing red cape, but do NOT reproduce any exact movie costume, logo, emblem, actor likeness, or copyrighted Superman design. The entire sequence is filmed from a very high overhead camera looking almost straight down at a huge outdoor film set. At first glance, the scene appears to show the superhero climbing vertically up the side of an enormous skyscraper. The illusion is created by a gigantic forced-perspective city set laid completely flat on the ground. The “skyscraper” is actually a highly detailed printed/constructed city facade positioned on the road, with perspective lines making it look vertical from the camera angle. Show the complete filmmaking environment: - the superhero lying/positioned on the fake building surface - multiple film cameras - camera operators - director - assistant directors - lighting equipment - cables - crew members - stunt coordinators - production assistants - spectators watching the unusual shoot - equipment cases and production vehicles around the perimeter The director is actively directing the superhero, gesturing toward different positions and calling out instructions. Crew members move around the set naturally, making the scene feel like genuine behind-the-scenes footage from a massive movie production. The superhero performs a dramatic “vertical climb” across the fake skyscraper. From the overhead perspective, the illusion looks incredibly convincing. Then create the viral reveal: The camera slowly pulls even farther upward, revealing the full scale of the setup and making it obvious that the entire skyscraper is actually lying flat on the ground. After the reveal, transition seamlessly from the overhead BTS shot into the “finished movie shot.” Suddenly the perspective changes and the same setup looks like an enormous vertical skyscraper, with the superhero appearing to climb the side of the building hundreds of feet above the city. Make the final cinematic shot look completely believable: dramatic sunlight, realistic atmospheric haze, detailed glass skyscrapers, physically accurate shadows, realistic cape movement, cinematic depth of field, subtle camera shake, premium blockbuster cinematography. The contrast between the ridiculous real-world setup and the spectacular final shot should be the main comedic hook. Style: ultra-realistic live-action, documentary BTS footage mixed with blockbuster superhero cinematography, natural human movement, realistic production equipment, physically believable lighting, detailed environments, authentic camera operators and crew. Camera progression: 1. Extreme overhead establishing shot. 2. Slow controlled aerial push-in. 3. Medium overhead shot showing the superhero and crew. 4. Wider reveal exposing the entire fake skyscraper. 5. Dramatic transition into the finished cinematic superhero shot. 6. Final heroic shot of the superhero climbing/flying beside the enormous skyscraper. No text, no subtitles, no logos, no watermarks, no recognizable actors, no exact recreation of an existing movie scene.show more

Johnn
42,663 görüntüleme • 8 gün önce
Seedance 2.0 on FlovaAI =================== Prompt: [Reference Identity Lock]... Image 1 is ONLY the main female protagonist. Her face, hairstyle, body type, and outfit must match Image 1 exactly and stay consistent for the entire video. Image 2 is ONLY a uniform reference. All four opponents wear the school uniform shown in Image 2. Never swap, merge, duplicate, or blend identities. The protagonist's identity comes ONLY from Image 1. The four opponents have NO reference images. They are defined by the text descriptions below. The four opponents must not resemble the protagonist, and they must not resemble each other. All five characters must remain clearly distinct and recognizable until the end. [Priority Order] 1. Preserve the protagonist's identity from Image 1. 2. Keep the four opponents visually distinct from her and from each other. 3. Maintain one continuous shot with no cuts. 4. Keep the classroom layout spatially consistent. 5. Make the action fast but readable and physically connected. 6. Keep the tone as a Korean school action drama, stylish but grounded. Korean school action drama classroom fight scene — 15 seconds, ONE CONTINUOUS SHOT, NO CUTS. A single uninterrupted handheld shot. No cuts, no scene transitions, no montage. The camera should feel handheld, with micro-jitters, slight rolling shutter, and raw unstable realism. The camera must physically travel through the same classroom space. Every transition must be motivated by camera movement, not editing. Whip pans are allowed, but they must not hide a cut. Do not teleport the camera or characters. The classroom layout and character positions must remain spatially consistent. Audio: No music. Only realistic school and classroom ambient sounds: old fluorescent light hum, distant hallway noise, ceiling fan, shoes scraping the floor, desks dragging, chair legs screeching, cloth friction, dull body impacts, and breathing that gradually becomes heavier. Breathing continues throughout the scene and keeps building. Lighting: Late afternoon in a Korean high school classroom. Mixed cool fluorescent light and warm sunlight through the windows. Dust floating in the sunlight. Soft fan shadows moving across desks and school uniforms. Main character: The Korean female high school student from Image 1, age 17–18. Cold, emotionless, calm, and intimidating. She barely speaks and does not scream during the fight. She remains composed from beginning to end. Her movements are efficient, explosive, and precise. Even if her frame is not large, she dominates through speed, timing, and accuracy. Main outfit: Exactly the outfit shown in Image 1. Do not change its colors, design, or details. Her jacket or outer layer is either removed and hanging on a chair, or worn in a slightly messy way. The action must be non-sexualized and combat-focused. Fabric movement, dust, sweat, wrinkles, and impact response should feel realistic. Opponent rules: Four Korean female high school students, all wearing the Hanlim Multi Art School uniform shown in Image 2. They have no reference images. Define them strictly by these descriptions and keep each one consistent: Opponent A: short black bob with straight bangs, medium build, round face. Opponent B: long straight hair tied in a high ponytail, tall and lean, sharp jawline. Opponent C: shoulder-length hair with side-swept bangs, slim build, narrow face. Opponent D: long wavy hair worn loose, slightly stocky and broad-shouldered. A, B, C, and D must each keep clearly different faces, hairstyles, body shapes, and silhouettes. They must not resemble the protagonist, and they must not resemble each other. No face duplication, no face merging, no identity confusion. Environment: An empty classroom at Hanlim Multi Art School, a Korean performing arts high school in Seoul. Green chalkboard, chalk tray, worn wooden desks, plastic chairs, classroom clock, class schedule poster, discipline/life-guidance posters, cleaning tools, blinds or curtains, wall study materials, and a slightly scuffed floor. Desks and chairs should react naturally to impacts, sliding, shaking, and collapsing when hit. Camera framing rules: Even during kicks, framing should stay around chest-level or eye-level. No low-angle shots under the skirt. Do not focus on legs, thighs, underwear, or fetish-like details. All action framing must prioritize faces, upper-body motion, impact, and spatial choreography. Continuous action and camera choreography: From 0 to 15 seconds, the fight continues without any cuts. The action should be stylish but readable, and every movement must be physically connected. 0–3s: The camera starts behind the protagonist at a slightly low handheld angle, drifting left through the classroom aisle. Opponent A grabs the protagonist's shoulder roughly and says in Korean: "야, 너 지금 뭐 하자는 거야?" The protagonist silently turns and lands one hard straight punch to A's face. At impact, use a very brief 15% slow motion: cheek ripple, dust particles, deep thud. A falls sideways into a desk. The camera dips slightly from the shock, then whip-pans right without cutting. 3–6s: Opponent B charges in from the right. The protagonist steps forward instead of retreating. A short body shot to the stomach. Immediate uppercut to the chin. Without pausing, she drives forward into a flying knee to B's chest. B is thrown backward across or into a desk. The camera follows the forward motion low, then rebounds upward with the impact. 6–9s: Opponent D attacks with two fast punches. The protagonist deflects both strikes with her arms, then flows into a turning backfist to D's face. As D staggers, she continues the same rotation into a spinning back elbow that lands hard on D's jaw or temple. D crashes sideways into two or three desks. The camera arcs around her shoulder and jitters slightly at each impact. No cuts. 9–12s: Opponent C rushes in from the chalkboard side. The protagonist clearly grabs C's collar with her left hand. C's face must be fully visible from the front and clearly different from the protagonist. The protagonist lands one short, hard punch to C's face, then immediately throws a powerful high kick or flying high kick into C's chest. The force sends C backward into the green chalkboard. The protagonist remains in the foreground and never touches the board. The protagonist's face should be side-profile or partially obscured. C's face should be clearly visible from the front at the moment of impact. Their faces must never overlap in frame. Use a very brief 20% slow motion at the chalkboard impact: chalk dust bursts outward, and C slides down the board. The camera pushes up with the impact, then tilts down as C slides. 12–15s: Through the chalk dust, the camera hard-pans right. D makes one final charge. The protagonist sidesteps and lands a tight uppercut to D's chin, followed immediately by a cross. D crashes into a row of desks, causing a chain reaction of collapsing desks and chairs. The camera drifts forward slowly. The protagonist adjusts her loose tie or ribbon and brushes chalk dust off her shoulder. Her expression stays cold and serious. She walks past the camera and exits the frame. Dust floats in the sunlight. Natural ending. =================== Made with Flova #FlovaAI #FlovaCPPshow more

TSUBAKI
19,167 görüntüleme • 1 ay önce
This workflow is perfect for creating short fashion-style cinematic... videos. I simplified the original prompts based on willie’s method, and the whole process is now much faster and more stable: 1. Generate a 3×3 keyframe grid (Nano Banana Pro only) Use this simple prompt: “In a 3x3 grid, show this character in different angles, keep the scene the same, random poses. This is far simpler and more efficient than my old prompts. You can generate multiple times and just pick the keyframes you like most. 2. Extract a high-res keyframe (Super stable trick) Take a screenshot of the keyframe you want from the 3×3 grid, send it back to Nano Banana Pro, and simply say: “Give me a high-resolution version.” This method is much more stable than relying on complex upscaling prompts. 3. Generate the video with Kling 2.5 Turbo Upload the first and last frames to Kling 2.5 Turbo and use this prompt: “The camera very slowly and smoothly lowers on a boom.” From my testing, Kling 2.5 Turbo offers the best balance of stability and cost — other models are either less consistent or noticeably more expensive. 4. Final speed adjustment with willie’s tool (Critical step) Use the tool built by willie to fine-tune the playback speed of each clip. This step is essential for getting that premium cinematic feel. I’ll drop the tool link in the comments.show more

underwood
22,223 görüntüleme • 9 ay önce
Messi thought he had this match under control... then... Yamal changed everything. Video for VivaReel prompt Create a fun, dynamic stop-motion style animated video in vibrant Lego bricks and minifigures aesthetic. The entire scene uses colorful plastic Lego construction with visible studs, bricks, plates, and minifigure details. All characters are Lego minifigures with classic yellow skin (or skin tones), printed faces, and detailed soccer uniforms made from Lego pieces. The environments are fully built from Lego: layered brick landscapes, trees, flowers, stadiums, mountains, and buildings with perfect Lego texture and lighting. The video tells a short humorous story , it shows a connected previous or parallel moment, Smooth transitions between scenes every ~1 second. Maintain consistent Lego papercraft diorama look but fully realized in 3D Lego bricks. Sequence: 1. Lamine Yamal Lego minifigure (red/blue Spain jersey #19, curly black hair) juggling and kicking a black/white Lego soccer ball on a winding Lego path through green fields, trees, flowers, and a distant stadium under a blue sky with sun and clouds. 2. Transition to Lionel Messi Lego minifigure (Argentina striped jersey #10, beard, tattoos) standing confidently with foot on ball in front of Lego Buenos Aires scenery (pink Casa Rosada, obelisk, mountains, Argentine flag). 3. Messi looks sad/frustrated, sitting on grass with blue tear streams, hands on face. Yamal minifigure runs past happily dribbling the ball. 4. Dramatic stadium scene: Yamal in Spain kit runs past two Argentina defenders and kicks the ball powerfully toward goal. Close-up of foot striking ball. 5. Yamal scores! Goalkeeper dives and misses. Yamal celebrates by lifting the golden Lego World Cup trophy high on the field with confetti raining down. Messi lies on the ground covering his face in defeat nearby. Use bright, cheerful Lego colors, dynamic camera angles (wide shots, action close-ups, low angles), smooth minifigure animations, and upbeat energetic feel. High detail Lego texture, cinematic lighting, 16:9 aspect ratio, 16 seconds duration." #happyhorse #vivareelshow more

Sharon Riley
76,294 görüntüleme • 1 ay önce
Three styles, three expressions captured in a perfectly designed... staircase. Photos & Video made with AI (Nano Banana) Edit in InShot My own Idea, Style & Design Prompted & edited by me 👉 Subscribe for exclusive videos + photo sets (Content you won’t see on the main feed) Prompt I used: "Create a photorealistic, high-detail lifestyle photograph of three young women posing together on a modern wooden staircase inside a stylish contemporary home. Do not reference or resemble any real celebrity or public figure. 1. Subjects, hair, skin, expressions & poses: Woman on the left: young woman with fair skin, light blonde hair styled in a neat high ponytail with a few soft strands framing her face. She has natural facial features, subtle makeup, and a relaxed, friendly expression. She sits comfortably on a wooden stair with one arm raised casually toward her hair, looking toward the camera with a confident but natural smile. Woman in the center: young woman with fair skin and blonde hair pulled into a sleek ponytail. She has softly defined eyebrows, natural makeup, and a calm, composed expression while looking slightly toward the side. She stands or sits one step higher than the others, creating a layered composition. Woman on the right: young woman with fair skin and long copper-red hair gathered into a ponytail, with a few loose strands around her face. She has subtle freckles, natural makeup, and a relaxed expression. She sits sideways on a lower stair with one hand resting naturally on the step while looking toward the camera. Keep all three women anatomically natural and proportionate, with realistic hands, facial symmetry, hair strands, and natural posture. Their interaction should feel like a casual group photograph between friends. 2. Clothing & accessories: Left woman wears a green-and-black horizontally striped sleeveless summer dress with a simple elegant design, paired with a delicate gold necklace and small earrings. Center woman wears a navy-and-blue striped sleeveless dress with a clean contemporary design and minimal jewelry. Right woman wears a burgundy-and-black striped sleeveless dress, with subtle jewelry and visible decorative tattoo artwork on her upper arm. Use realistic fabric texture, stitching, folds, and natural draping. Keep the styling fashionable but tasteful and suitable for a casual lifestyle photograph. 3. Environment & lighting: Set the scene inside a bright, modern multi-level home with a distinctive wooden staircase, white structural beams, thin metal cable railings, and warm wooden steps. Include contemporary architectural details, glass panels, neutral walls, minimalist furniture, and subtle decorative elements in the background. Large windows allow soft daylight to enter the room. Use warm ambient interior illumination combined with natural daylight for a welcoming atmosphere. Create realistic shadows and gentle highlights across the subjects and staircase without excessive contrast. Background should have moderate depth-of-field blur while retaining enough architectural detail to establish the location. 4. Camera & visual style: Photorealistic editorial lifestyle photography. Shot on a 50mm full-frame lens, approximately f/2.8, with natural perspective and subtle background separation. Eye-level camera positioned slightly below the group to emphasize the staircase architecture while keeping faces clearly visible. Vertical portrait composition, approximately 4:5 aspect ratio. Natural skin texture, realistic hair detail, accurate fabric texture, physically realistic lighting, sharp facial details, and authentic photographic depth. Warm cinematic color grading with balanced skin tones, subtle contrast, gentle highlights, and natural saturation. High dynamic range, professional indoor photography, crisp focus on all three subjects, realistic depth of field, ultra-detailed, polished but not artificially airbrushed."show more

J⭕DIE
13,947 görüntüleme • 24 gün önce
WATCH THIS VIDEO CAREFULLY. FORENSIC ANALYSIS OF A VIDEO... CURRENTLY BEING CIRCULATED AND SPREAD ON ARAB TELEGRAM CHANNELS (Mor Edge Insight in conjunction with GAZAWOOD - The Pallywood Saga - BACKUP - July 6) What you are about to see is raw footage of an active arrest operation and genuine footage. This clip is currently circulating on Palestinian Telegram channels and is being prepared for wider distribution on X. It follows a familiar pattern of real footage with heavy manipulation and inauthentic audio to create a perception and narrative that doesn’t exist and is not what the footage actually shows. Here is the step-by-step forensic breakdown. The audio track contains multiple sharp “gunshots.” However, frame-by-frame examination shows no muzzle flashes at any point, even in bright daylight where unsuppressed firearms would produce clear, visible bursts. There is also no visible recoil or weapon movement on the individuals holding rifles. The barrels show no suppressors, yet the sounds are relatively clean “pops” rather than the overwhelming cracks expected from unsuppressed fire at that range. The audio of the shots fired are more reminiscent of a children’s toy than a real gunshot. More critically, the visual action is happening at a clear distance across the road, at a distance of an estimated 60-100m away from the camera, yet the gunshots and shouting sound as if recorded right next to the camera. Real distant gunfire would be thinner, more muffled, and accompanied by environmental echoes. This audio was added in post-production. How distance was determined: The white car in the immediate foreground (partially visible on the left) is only 5–10 meters away. The road width and the position of the parked vehicles and people with guns put the core action clearly in the mid-ground, across the full width of the street and shoulder. Reference objects: Standard car lengths (4.5–5m), average adult height (1.7m), and the spacing of streetlights/power poles all support a distance in that 60–100 meter range for the shooters and the SUV. The black SUV drives a noticeable distance across the frame without appearing overly large or close, further confirming it’s not right next to the camera. This distance makes the audio mismatch even more obvious. Real gunfire at 60–100 meters would sound significantly more distant and muted, with clear delay and environmental filtering. The overlaid “cracks” sound like they were recorded (or synthesized) much closer. Summary 1. Real gunshots, especially in an open outdoor environment like this, produce a sharp initial crack (supersonic bullet) followed by a broader report/echo, with significant low-frequency rumble, reverberation off the ground/cars/objects, and environmental decay. These sound more like clean “pop/crack” samples layered on top. 2. They lack the natural variations in volume, timing, or distortion you’d expect from actual firearms in a real chaotic scene (muzzle blast, echoes, distance differences) even with silencers which from that distance you wouldn’t even hear. They feel “pasted in” during editing. 3. The overall audio mix (ambient road noise, car sounds, voices) doesn’t interact naturally with the “shots”, there is no proper masking, reverb bleed, or mic overload you’d get from real loud events captured on the same recording device. Always examine the audio against the visuals, check for continuity errors, and watch how people actually behave when they think no one is watching the performance. Share if you value this kind of detailed verification.show more

Mor Edge Insight
23,394 görüntüleme • 2 ay önce
Only used the character sheet quoted below. I gave... up generating storyboards as it kept redrawing the helmet, even when a reference sheet was given exclusively for it. Seedance likes structured and concise prompts so I gave this format a shot. Text to video prompt: Use the attached character sheet as the STRICT character and helmet reference. Create a 15-second cinematic stylized 3D animation. IMPORTANT: The character sheet controls the final design exactly. Do not redesign the helmet. Maintain the exact silhouette: two massive gold crescents curving inward, large centered red sun disk, rounded gold helmet cap, front red jewel, large round ear ornaments. STYLE: Cute dry exaggerated comedy. Soft cinematic lighting. Stylized 3D animation. Grounded acting and believable weight. No chibi proportions. No anime combat energy. SCENE: Early morning inside an Egyptian-inspired palace bedroom. Warm sunrise through curtains. Simple elegant room with bed, side table, mirror, doorway. SHOT FLOW: 1. She sleeps in bed while the oversized helmet rests nearby on a table. 2. She slowly wakes up, notices the helmet, and immediately looks exhausted and annoyed. 3. Dramatic close-up of the helmet sitting silently like a daily burden. 4. She walks toward it with sleepy acceptance. 5. She grabs the helmet with both hands and struggles lifting it because it is extremely heavy. 6. She raises it over her head while wobbling from the weight. 7. She lowers the helmet onto her head ONCE. It lands crooked and squishes her hair awkwardly. 8. Without removing it, she aggressively twists and adjusts it into the correct position while visibly frustrated. 9. She grabs a tiny morning drink and walks out into bright morning sunlight still looking dead inside. ANIMATION PRIORITIES: subtle facial acting, comedic pauses, helmet heaviness, small body balance corrections, secondary motion in hair, cloth, jewelry, and sash, clean cinematic staging. CAMERA: slow cinematic push-ins, medium acting shots, clean wide shots, subtle handheld wobble during struggle moments. AVOID: fight choreography, magic, speed lines, hyperactive motion, slapstick chaos, helmet redesigns, extra accessories, anime exaggeration.show more

Glitter Gal
13,013 görüntüleme • 3 ay önce
AI Is Moving Beyond “Generating Videos” — Toward “Generating... Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:show more

雪踏乌云
113,347 görüntüleme • 1 ay önce