Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

i created this WWII documentary just with AI video models no After Effects no Remotion or complex setups no $1000 spent on an editor it took Calliope 30 seconds (yes, 30 seconds) to animate all the clips in parallel here's how: 1- reuse the image + video prompts from...

102,232 görüntüleme • 18 gün önce •via X (Twitter)

1 Yorum

Brjan | AI Builder profil fotoğrafı
Brjan | AI Builder18 gün önce

30 seconds is impressive, but the results can vary widely based on prompt quality

Benzer Videolar

This is probably the most complex workflow I’ve ever built, only with open-source tools. It took my 4 days. It takes four inputs: author, title, and style; and generates a full visual animated story in one click in ComfyUI . I worked on it for four days. There are still some bugs, but here’s the first preview. Here’s a quick breakdown: - The four inputs are sent to LLMs with precise instructions to generate: first, prompts for images and image modifications; second, prompts for animations; third, prompts for generating music. - All voices are generated from the text and timed precisely, as they determine the length of each animation segment. - The first image and video are generated to serve as the title, but also as the guide for all other images created for the video. - Titles and subtitles are also added automatically in Comfy. - I also developed a lot of custom nodes for minor frame calculations, mostly to match audio and video. - The full system is a large loop that, for each line of text, generates an image and then a video from that image. The loop was the hardest part to build in this workflow, so it can process either a 20-second video or a 2-minute video with the same input. - There are multiple combinations of LLMs that try to understand the text in the best way to provide the best prompts for images and video. - The final video is assembled entirely within ComfyUI. - The music is generated based on the LLM output and matches the exact timing of the full animation. - Done! For reference, this workflow uses a lot of models and only works on an RTX 6000 Pro with plenty of RAM. My goal is not to replace humans, as I’ll try to explain later, this workflow is highly controlled and can be adapted or reworked at any point by real artists! My aim was to create a tool that can animate text in one go, allowing the AI some freedom while keeping a strict flow. I don’t know yet how I’ll share this workflow with people, I still need to polish it properly, but maybe through Patreon. Anyway, I hope you enjoy my research, and let’s always keep pushing further! :)

Lovis Odin

58,841 görüntüleme • 1 yıl önce

Claude Code can ship a 45-second animated explainer ad in 30 minutes. No video editor needed, just CC + skills. Here's how I made this video for Soteri Skin 👇 1. /plan Concept Brief (Claude Code) I handwrite a concept brief, then chat with the agent to iterate on it. The agent gathers any raw materials we might need - context about the brand, product images, end card, etc. The concept brief details the concept, characters, visual style, script, etc 2. /prepare a moodboard (CC + GPT Image 2 + ElevenLabs) After reviewing the script, generate: - character reference images - voiceover samples for the characters / narrator - the storyboard (scene by scene grid) - a few keyframe scenes 3. /generate Keyframes for each scene (CC uses Nano Banana or GPT Image 2) Uses the character references from the previous step to generate keyframes for each scene. I probably should have done a round of iteration at this step – there's some character drift and the pH meter representation could have been better. 4. /animate Keyframe → Animated Clip (CC uses Fal Seedance) Generate 2-4 representative scenes first to see a preview. If it looks good, then generate everything. 5. /stitch (CC + ffmpeg + ElevenLabs) - Stitch clips together with hard cut - Add a music score + SFX - Sync clips to the VO - Add captions - Review and edit timing / pacing issues 6. /watch the final cut and review it - as a video editor for technical errors (mismatched voiceover and visuals, AI hallucinations, etc) - as a viewer (ICP). I delegate most of the review to the agent because it catches more things and keeps me out of the loop as much as possible. It also fixes any issues found in the review. That's it. This video took me 30 minutes because I have already created skills for everything I described above. Some day, this will be < 5 minutes. I just review and chat to provide direction and feedback. The skills do all the technical work. 7. /learn Extracts learnings and updates the skills. This final step is really important. It turns this process into a closed loop system that makes the next video much easier to create because all the learnings from the human-in-the-loop process get encoded into code. Skills are code too. If you want access to the skill, drop a comment, and I'll DM it to you (must be following). If you want to make AI video ads like this, DM me.

Shiv

11,760 görüntüleme • 4 ay önce