currently experimenting with WAN 2.6 I2V on GMI Cloud... in this test, I’m comparing two audio workflows and honestly both perform really well. one scene uses audio generated directly from the prompt, while the other uses manually uploaded audio taken from the film 300. visually, both deliver strong motion and solid performance. however, the version with audio coming straight from the prompt feels slightly more refined, camera movement is smoother, transitions flow more naturally and the sync between voice, facial motion, and pacing feels more cohesive. lip sync, especially for Chinese dialogue also comes across a bit cleaner. you can choose single shot for a clean, focused moment or multi shot if you want more cinematic transitions, even when working from just one reference image. one important note: always turn on prompt extension. it makes a noticeable difference in how well the model understands motion, transitions and overall scene flow. both audio approaches are totally usable, but if you’re building dialogue driven or cinematic scenes, starting with audio from the prompt gives WAN 2.6 a bit more context to work with. I’ll be pushing this further with more dynamic camera movement and transitions next. more experiments coming soon ✨ Wanshow more

DStudioproject
96,950 Aufrufe • vor 7 Monaten
Seedance 2.0 from China will be the SOTA This... is AI We are cooked. • Native multi-shot storytelling from a single prompt (no more stitching scenes) • Phoneme-level lip-sync in 8+ languages • 30% faster generation than v1 via RayFlow optimization • 1080p cinematic quality, maintains motion consistency across shots • Audio-video joint generation = no separate audio pipeline • Multi-shot coherence = you could generate a full short film scene in one promptshow more

Dorksense
388,807 Aufrufe • vor 5 Monaten
This AI just turned me into a film director…... No editing skills. No timeline headaches. Just one prompt. This is Seedance 2.0 🎬 You can literally combine: → Text → Images → Videos → Audio And it understands everything. Even crazier? You can control it like this: Image → character Video → camera movement audio1 → music/voice It doesn’t just generate clips… It builds full cinematic scenes with: → Consistent characters → Smooth transitions → Realistic motion → Built-in lip sync Basically… From a single prompt → you get a multi-shot story. Not AI video. AI filmmaking. Go try it before everyone catches on 👇show more

Kshitij Mishra | AI & Tech
60,393 Aufrufe • vor 3 Monaten
GPT Image 2 + Seedance 2.0 Prompt Share Created... on mitte.ai I didn't use a character sheet for this generation. I directly used the character images I created in Midjourney. Since the visual style transfers into the video surprisingly well, it's actually a really effective for style transfer. This time I also added a bit more detail to the Seedance prompt itself. You can definitely achieve similar results without storyboards too, they're not mandatory but I think they're one of the best ways to previsualize scenes, pacing and even camera angles before generation. Also, this storyboard prompt is still a bit long. I'm currently experimenting with more compact version of it too. You can check the prompts below.show more

Kōda
38,856 Aufrufe • vor 1 Monat
GPT Image 2 + Seedance 2.0 Prompt Share Mei... Lin's Elemental Kung Fu Performance Created on mitte.ai I'm experimenting with Laban right now. I think it helps make movements feel smoother and more expressive, but I still need to do more tests. For this one, I only added a small Laban section to the Seedance prompt. Laban is a movement analysis system that describes motion through weight, time, space and flow to shape expressive body language and performance. You can find the prompts below.show more

Kōda
84,664 Aufrufe • vor 2 Monaten
🎬 Motion Control has arrived! Take full control of... your AI videos with 12 dynamic camera shots — from smooth Dolly moves to dramatic Cranes and VFX like Explosions and Disintegration. Perfect for adding cinematic flare, stock footage, product ads, or just experimenting with storytelling. ✨ How it works: 1️⃣ Head to the new Video creation tool 2️⃣ Enter your prompt 3️⃣ Pick your camera shot from the Motion Control panel 4️⃣ (Optional) Add a Style or inspiration image as a start frame 💡Want extra consistency? Train an Element, generate your key stills, and animate them for a fully guided scene. More features are coming soon — including End Frames 👀 So stay tuned, and let us know what you create! 🎥 Lights, camera... Motion! Try it now on the Video page 👉🏻show more

Leonardo.Ai
984,421 Aufrufe • vor 1 Jahr
🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns... 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.show more

Rohan Paul
63,859 Aufrufe • vor 10 Monaten
Seedance 2.0 is insane... AI filmmaking is no longer... locked behind complex workflows, expensive tools, or regional limits. SJinn Agent now supports both Seedance 2.0 Pro and Seedance 2.0 Fast, giving creators a faster way to generate cinematic videos with more control over the final result. You can add image, video, and audio references to guide the direction, motion, style, and feeling of your videos, making the process more like directing And the best part: they’re offering 40% off, so this is probably the easiest time to test what high-level AI video creation can actually look like Prompt in first comment:show more

Amira Zairi
55,860 Aufrufe • vor 2 Monaten
🚀 Update Next Scene V2 only 10 days after... last version, now live on Hugging Face 👉 🎬 A LoRA made for Qwen Image Edit 2509 that lets you create seamless cinematic “next shots” — keeping the same characters, lighting, and mood. I trained this new version on thousands of paired cinematic shots to make scene transitions smoother, more emotional, and real. 🧠 What’s new: • Much stronger consistency across shots • Better lighting and character preservation • Smoother transitions and framing logic • No more black bar artifacts Built for storytellers using ComfyUI or any diffusers pipeline. Just use “Next Scene:” and describe what happens next , the model keeps everything coherent. 🧩 Try it directly in ComfyUI, or check the thread to launch it on fal . Open-source, no restrictions, made for filmmakers, animators, and dreamers. ComfyUI #AIcinema #LoRA #Flux #Qwen #ComfyUI #AIart #GenerativeVideo you can test on comfyui or to try on you can go here : and use my lora link : start your prompt with "Next Scene:" and lets go !!show more

Lovis Odin
43,276 Aufrufe • vor 9 Monaten
1/ We've all been aware of the hype surrounding... Dreamina Seedance 2.0, and I finally got early access to the tool. And wow, I'm blown away. This is by far the best. It is starting to make AI video feel less like "generate a clip" and more like "direct a scene." What stood out to me is the level of control: camera motion, pacing, visual consistency, and the ability to build from multiple references inside one workflow. Some of the prompt directions that feel especially strong: - a busy modern city square during daytime. Suddenly, time freezes completely - a single continuous camera movement through a natural landscape that transitions through all four seasons in one shot - an underwater bioluminescent city waking up at dawn The big shift is this: One Prompt, Viral Remade. Edit Videos as Easy as Editing Photos. Dreamina Seedance 2.0 feels like a real step toward AI-native directing rather than just AI generation. Here are some examples 🧵:show more

Chubby♨️
75,394 Aufrufe • vor 3 Monaten
Opus 4.6 vs GPT-5.4 (4/9) prompt: Build a production-quality... 3D flight-tracking web app using React + Vite + Three.js (react-three-fiber + drei) that visualizes live OpenSky aircraft data on a rotatable 3D Earth, with real-time plane motion, smooth interpolation, altitude-accurate positioning, and polished lighting/post-processing. Both models did really well on this one and honestly I’m impressed with both. GPT-5.4 had the nicer post-processing out of the box. I really liked the subtle light shimmer on the airplanes when rotating the planet, and the camera work when clicking a plane felt better overall. Opus 4.6 though had a few details I liked more. It automatically went and found a much nicer Earth texture on GitHub, while with GPT-5.4 I had to reprompt it to go look for a better one. I also preferred Opus’s plane model overall, it just looked more polished, whereas GPT-5.4’s plane asset looked a bit funny. One thing I noticed with GPT-5.4 is that when you click into the plane, the camera sometimes clips through the planet, which breaks the effect a bit. Opus handled that part more cleanly. Overall this felt like a strong result from both, just with different strengths. GPT-5.4 felt better on presentation and post-processing, while Opus had better asset choices and a more premium-looking Earth/plane combo.show more

Dev Ed
279,565 Aufrufe • vor 4 Monaten
WATCH THIS VIDEO CAREFULLY. FORENSIC ANALYSIS OF A VIDEO... CURRENTLY BEING CIRCULATED AND SPREAD ON ARAB TELEGRAM CHANNELS (Mor Edge Insight in conjunction with GAZAWOOD - The Pallywood Saga - BACKUP - July 6) What you are about to see is raw footage of an active arrest operation and genuine footage. This clip is currently circulating on Palestinian Telegram channels and is being prepared for wider distribution on X. It follows a familiar pattern of real footage with heavy manipulation and inauthentic audio to create a perception and narrative that doesn’t exist and is not what the footage actually shows. Here is the step-by-step forensic breakdown. The audio track contains multiple sharp “gunshots.” However, frame-by-frame examination shows no muzzle flashes at any point, even in bright daylight where unsuppressed firearms would produce clear, visible bursts. There is also no visible recoil or weapon movement on the individuals holding rifles. The barrels show no suppressors, yet the sounds are relatively clean “pops” rather than the overwhelming cracks expected from unsuppressed fire at that range. The audio of the shots fired are more reminiscent of a children’s toy than a real gunshot. More critically, the visual action is happening at a clear distance across the road, at a distance of an estimated 60-100m away from the camera, yet the gunshots and shouting sound as if recorded right next to the camera. Real distant gunfire would be thinner, more muffled, and accompanied by environmental echoes. This audio was added in post-production. How distance was determined: The white car in the immediate foreground (partially visible on the left) is only 5–10 meters away. The road width and the position of the parked vehicles and people with guns put the core action clearly in the mid-ground, across the full width of the street and shoulder. Reference objects: Standard car lengths (4.5–5m), average adult height (1.7m), and the spacing of streetlights/power poles all support a distance in that 60–100 meter range for the shooters and the SUV. The black SUV drives a noticeable distance across the frame without appearing overly large or close, further confirming it’s not right next to the camera. This distance makes the audio mismatch even more obvious. Real gunfire at 60–100 meters would sound significantly more distant and muted, with clear delay and environmental filtering. The overlaid “cracks” sound like they were recorded (or synthesized) much closer. Summary 1. Real gunshots, especially in an open outdoor environment like this, produce a sharp initial crack (supersonic bullet) followed by a broader report/echo, with significant low-frequency rumble, reverberation off the ground/cars/objects, and environmental decay. These sound more like clean “pop/crack” samples layered on top. 2. They lack the natural variations in volume, timing, or distortion you’d expect from actual firearms in a real chaotic scene (muzzle blast, echoes, distance differences) even with silencers which from that distance you wouldn’t even hear. They feel “pasted in” during editing. 3. The overall audio mix (ambient road noise, car sounds, voices) doesn’t interact naturally with the “shots”, there is no proper masking, reverb bleed, or mic overload you’d get from real loud events captured on the same recording device. Always examine the audio against the visuals, check for continuity errors, and watch how people actually behave when they think no one is watching the performance. Share if you value this kind of detailed verification.show more

Mor Edge Insight
23,103 Aufrufe • vor 19 Tagen
Kling AI 3.0 is here And it’s a serious... step forward for AI video. This update isn’t about small tweaks. It’s about refinement, realism, and making AI video feel ready for real-world use. What stands out with Kling 3.0? Stronger visual quality. Movements feel smoother. Lighting looks more natural. Scenes feel intentional instead of generated. Better prompt understanding. You can describe complex scenes, moods, or camera directions, and Kling 3.0 interprets them with much more accuracy. Less trial and error. More usable results. Improved motion and consistency. Characters stay consistent. Objects behave logically. Shots flow more naturally from start to finish. Greater creative control. From product-style visuals to cinematic storytelling, Kling 3.0 gives creators the flexibility to move beyond simple clips and into structured, high-quality sequences. The difference is subtle until you see it, and then it’s obvious. AI video is evolving quickly, and Kling AI 3.0 shows how far the technology has come. It’s not just about generating video anymore. It’s about generating video that’s usable, polished, and ready for campaigns, storytelling, and branded content. If you’re building with AI, this is one of the tools worth paying attention to, best AI video model for pro & commercial productionshow more

Future Stacked
183,095 Aufrufe • vor 5 Monaten
Introducing Muse Image and Muse Video, the first media... generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities to image generation and integrates with Muse Spark. You can try Muse Image in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Today we’re also previewing Muse Video, which is built upon the same pretraining base as Muse Image to deliver exceptional visual fidelity with native audio support. Learn more about both models:show more

AI at Meta
830,454 Aufrufe • vor 17 Tagen
gemini omniflash is actually f*cking cracked. you can animate/edit... any video with a text prompt. character swaps, object transforms, full environment changes without regenerating/rotoscoping. everyone using AI to to animate and edit videos right now hits the same wall. the clip comes out 90% right and you regenerate from scratch hoping the 10% fixes itself. it never does. the fix is using your video as the input. omniflash edits what's already there instead of rolling the dice again. here's what's in the system: > the two-layer premiere trick: generate the same shot twice (one with background removed), stack them, cut at one frame, instant scene change > character swap with a single reference image (plus the one line you need or the model keeps the original's features) > object transforms that leave the rest of the frame untouched: stone into glowing sphere, candles into flowers > style transfer from an image reference instead of text, way more accurate > why stacking edits in one prompt breaks everything and the exact step order that doesn't > the audio limitation nobody mentions and how to work around it i packaged every prompt, the edit sequence, and the premiere layering setup. RT + reply "OMNI" and i'll send it over.show more

Sulfur
36,015 Aufrufe • vor 22 Tagen
I tried Hailuo AI (MiniMax) to see how it... handles real content creation. The workflow is simple. You just write a prompt or drop in an image, and it turns that into a dynamic video with motion, framing, and scene depth. No timeline to manage. No editing setup. No back and forth. What stood out to me: • Text to video and image to video both feel smooth. • It handles motion, camera angles, and flow on its own. • Output is fast, usually within seconds. • Works well for reels, quick ads, storytelling, and idea testing. It removes the hardest part: starting from scratch and turns your ideas into content in minutes. Instead of thinking, “How do I make this video?” You start with, “What do I want to create?” That shift alone makes it worth exploring. Try it here: #Hailuoshow more

Manish Kumar Shah
27,662 Aufrufe • vor 3 Monaten
Memories are the previous learnings and context. Developers build... up state and context as they do work, which is what makes a more tenured developer more effective than a brand new one with a similar skillset. It would be both inefficient and painful to relearn information from scratch around code structure, architecture, etc each time you needed to do a new task. Memories solve for this. As you do work with Cascade, it can automatically choose to “remember” pieces of information that it learns as Memories, and for any later work, it can choose to pull from this memory bank instead of trying to relearn that information from scratch. You also can manually prompt Cascade to remember parts of conversations as Memories and can manually go in and edit Memories post-fact. Here’s a developer asking Cascade to save some knowledge as a Memory:show more

Windsurf
12,356 Aufrufe • vor 1 Jahr
✨ I updated Photo AI's video model to Kling... AI 2.6 now By default it also generates audio (like Veo 3 does) and if you prompt for it it also produces voice, and it's pretty realistic too I'd say the cadence of the voice can be improved to be a bit more real, but it's definitely getting there It's just one click, just generate an AI photo first then hover over it and press [ 📼 Make video ] It's 3x more expensive, but I sell it below cost (5 credits) so that you all sign up to Photo AI and generate your UGC influencers with me 😊 Here's my 100% AI influencer telling what she thinks of the European Commissionshow more

@levelsio
323,674 Aufrufe • vor 7 Monaten
Honestly, I hate that I even have to say... this, but seeing people use one single fancam to claim Sana “can’t dance” is genuinely ridiculous. Y’all, for everyone dragging Sana because of one video from the 73rd THIS IS FOR tour show, I think it’s important to look at the full context before judging her dancing ability. First of all, Sana had been dealing with a cold, flu, and cough for over a month at that point, which can definitely affect stamina and performance consistency during a 2–3 hour concert. There are also several factors that can make the dancing look less smooth in that particular fancam: 1. Outfit If you’ve watched other fancams from this black outfit era, Sana was adjusting her outfit quite often on stage because it seemed to shift, slip, or sit unevenly at times. An uncomfortable outfit can naturally affect movement and make a performer more cautious. 2. Camera tracking The camera appears to follow her movements slightly late. Even a small delay can make smooth transitions look jerky or abrupt when viewed on video. 3. Awkward zoom distance The fancam isn’t zoomed out enough to show the full choreography, but it’s also not close enough to focus on facial expressions. Because of that, viewers end up focusing mostly on body transitions and posture, which can make movements look harsher than they actually are. 4. Phone camera limitations High-energy movements recorded on a phone can suffer from motion blur, stabilization issues, and frame-rate limitations. Sharp movements that look clean in person can appear choppy or less fluid on video. 5.Stamina This was already the 73rd show of the tour. Performing the same demanding choreography for dozens of concerts while dealing with illness can affect anyone’s energy level. And if this one clip is enough to convince you that Sana “can’t dance,” then I encourage you to watch other performances of Right Hand Girl from different angles and different outfits. Looking at a performer’s overall body of work will always give a more accurate picture than judging them from a single fancam. One fancam does not erase years of consistently solid performances.show more

puteri🍉
38,216 Aufrufe • vor 1 Monat
I've seen a lot of animatics lately that are... super detailed, essentially viewport previews of the final shot. But an animatic at it's core doesn't need to be anything fancy. Its main purpose is to plan and test the timing, camera angles, movement, and overall composition. With animatics you have to keep in mind the following: Pre-visualization: It allows to see a basic version of the animation or scene before committing to detailed work. Timing and pacing: An animatic helps identify how long each scene or shot should last, ensuring that the timing feels right. Planning: It helps with layout, camera angles, and transitions. By visualizing the shots, the team can ensure that the framing and overall design of the scenes work well in 3D space. Efficiency: It allows the team to test and fix any potential issues early on, like awkward movements or awkward pacing, before spending time on high-quality rendering or complex animation. In short, an animatic helps in conceptualizing the final animation by giving a low-res, rough version of the scenes, which guides the entire production process in terms of visual design, timing, and storytelling.show more

Voxyde
22,244 Aufrufe • vor 1 Jahr
Currently testing multi-shots with Kling AI 3.0 to see... how it holds up. All of this was generated from a single image. This is also the raw audio (definitely an improvement over the delivery of dialogue). Consistency is honestly very impressive. Poor Barry, he just loves honey 🍯 Prompt: Shot 1 (3s): A cinematic shot of a British detective interrogating a bear in handcuffs. Medium over the shoulder shot as the British man says to the bear "Come on Barry! Why do you keep doing this? This is the third time in 3 months now! Shot 2 (4s): Extreme closeup focused on the bears face, it's expression is shame, it just looks down and says "I can't help it, I'm addicted, I really love honey" Shot 3 (4s): Extreme closeup of the Britsh detective with a disappointed expression. Shot 4 (4s): Extreme closeup of the bears T-shirt that says I love honey and then cut back to the detectives face looking disappointed and shaking his head. Off camera you can hear the bear say "I'm sorry, I won't do it again"show more

Travis Davids
11,275 Aufrufe • vor 5 Monaten