Загрузка видео...

Не удалось загрузить видео

На главную

[Most robots react. This one thinks a step ahead.] Ant Group's Robbyant just published LingBot-VA 2.0 — a video-action foundation model built from scratch for robot control, not fine-tuned from a video generator. The usual approach takes a video generator made for content creation and bolts a robot policy...

196,499 просмотров • 18 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Furniture assembly is the task everyone name-drops and nobody actually attempts at real scale. Every demo I have seen is a scaled down IKEA leg or a single arm on a toy chair. This paper does it properly, real scale, bimanual, up to 7 subtasks and 1,550 control steps per episode, and it is validated on a real Kinova Gen3, not just in sim. That real-robot number is the one that matters: only a 16 percent drop on the hardest task going from simulation to hardware. That is a small enough gap to take seriously, and it did not happen by accident. They built a VR teleoperation rig specifically for coordinated dual-arm collection, because generic single-arm teleop setups do not capture the coordination real assembly needs, and the model predicts a continuous progress signal alongside the action chunk rather than a discrete subtask label, letting it auto-transition and catch drift before it compounds into total failure. The simulation ablation is what got them there, 48 to 80 percent over baselines, with another 21 points from their perception and control design study alone, but that is groundwork, not the headline. Watch the video, there is a clip of the robot misgrasping the seat panel, reopening the gripper, and regrasping on its own. That is not scripted recovery behaviour, it emerged from training, and it emerged on hardware. Excellent work from the team from Mitsubishi Electric Research Laboratories, with Oxford and UNC Chapel Hill Clinical Laboratory Science. Video and project page in comments. #Robotics #Manipulation #VLA

Stephen James

14,952 просмотров • 21 дней назад

📖THE STEP MOST CREATORS SKIP IS WHY THEIR AI ANIMATION LOOKS INCONSISTENT Consistency across clips doesn't come from prompting — it comes from the reference image. The pipeline, step by step: ▪ Start with ChatGPT Image 2 — generate a full character design sheet first, not just a single frame. Multiple angles, expressions, and outfit variations in one image keeps the character consistent across every scene ▪ Build a storyboard inside ChatGPT Image 2 as well — define each shot, camera angle, action, and mood before touching Seedance at all. This is the step most people skip and it's the reason clips look disconnected ▪ Define a color palette and lighting mood early — golden afternoon light, soft warm tones, dramatic shadows. Lock those values and repeat them across every prompt ▪ Take each storyboard frame into Seedance 2.0 as the reference image — one frame becomes one clip ▪ Write the Seedance prompt around the character action, not the scene description. The scene is already in the image. The prompt handles motion, camera behavior, and timing ▪ Keep clip duration between 4-6 seconds per shot — shorter clips give more control over pacing and reduce motion drift on character faces ▪ Match camera movement type across consecutive clips — if one shot dollies in, the next should hold or pull back, not dolly again The consistency across these frames comes from the character design sheet, not from luck. Seedance reads the reference image and the prompt together — if the reference is detailed enough, the output stays on-model. This video was created by ALOKXMEHTA 📥 tomorrow: the exact ChatGPT Image 2 prompt structure used to generate a multi-angle character design sheet like this one 🔖One article covers the entire workflow — it is pinned below, do not scroll past it.

Zentrix⌚️

12,846 просмотров • 28 дней назад

This week is already so hot. 🔥 Massive release from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."

Rohan Paul

12,761 просмотров • 6 месяцев назад

AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:

雪踏乌云

112,114 просмотров • 13 дней назад

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 просмотров • 6 месяцев назад

Figure 03 just finished an 8-hour work livestream, imperfect, but already good enough to replace a lot of repetitive warehouse labor. 🤖 Brett Adcock put a team of F.03 robots on a factory-style package sorting task for a full shift. The job was simple and brutal: detect the barcode, pick the package, flip it label-side down, place it on the conveyor, repeat. Soft poly bags, rigid boxes, moving belts, messy orientations. That is exactly the kind of boring physical work factories pay humans to do all day. Early in the stream, the system handled 230 packages in 10 minutes. That is roughly 2.6 seconds per item — already in human-speed territory for this narrow workflow. The more important part: it was not one robot pretending to work all day. It was a team of Figure 03 robots keeping the line running. When one robot ran low on battery, it left the station and another robot stepped in. That is the real factory signal: not just autonomy, but shift continuity. F.03 is rated for about 5 hours of runtime, so the 8-hour result depends on fleet orchestration, charging, and handoff. That matters more than a single clean demo. The stream was not perfect. There were pauses, hesitations, missed orientations, and small recovery moments. Good. A perfect short clip hides failure. An 8-hour livestream exposes the parts that actually matter: endurance, recovery, throughput, and whether the robot can stay useful after the novelty wears off. Figure says this was fully autonomous on Helix-02, with zero human intervention. For logistics and manufacturing, that is the threshold worth watching. Not “can it do one impressive task?” Can it keep doing the boring task for an entire shift? Figure is not showing a general human replacement yet. But for structured, repetitive factory work, the gap just got much smaller. The timing is also interesting: Figure says BotQ has already delivered 350+ F.03 units and reached a 1 robot/hour production cadence. And F.04 is now in full design lock, with parts starting to ship. The next test is obvious. 8 hours was the proof of endurance. 24/7 is the proof of labor economics.

RoboHub🤖

16,818 просмотров • 2 месяцев назад

BOOM! Humanoid Robots Just Performed Surgery for the First Time! REAL VIDEO! In a groundbreaking preclinical breakthrough, researchers at UC San Diego have achieved what many thought was years away: teleoperated humanoid robots successfully completing live surgeries. Published in Nature, the study marks the world’s first use of humanoid robots for in-vivo laparoscopic procedures on large animals (pigs). Two separate surgeries were completed: Key Details •. Procedure: Laparoscopic gallbladder removal (cholecystectomy) •. Team 1: Human surgeon + one humanoid robot (the robot performed core tasks while the human assisted) •. Team 2: Two humanoid robots working together with no human at the operating table •. Robots: Custom “Surgie” humanoids (~5 ft tall, ~60 lbs) using standard surgical tools •. Control: Fully teleoperated by surgeons (remote human control, not autonomous) •. Significance: First demonstration of humanoid robots handling real surgical workflows in a live setting, proving compatibility with existing OR tools and spaces This proof shows humanoid robots could one day help address surgeon shortages, enable remote procedures in rural areas, battlefields, or even space all at a fraction of the cost and space of traditional surgical robots like da Vinci. Read the full publication here: Project page with video: The future of surgery just got a whole lot more interesting. And medical cost for the first time in decades will be scheduled to go down, much further down.

Brian Roemmele

107,553 просмотров • 20 дней назад

Beauty ads just changed forever. Free Claude Opus 4.8 + GPT Image 2 + Seedance 2.0 workflow to spin up 100s of video ads. No studio, no model, no macro lens, no shoot day. Here's what nobody in beauty marketing wants to say out loud. That glossy lip shot. The droplet hitting the surface in slow motion. The whip-pan into the next scene. The crystalline product splash. All the stuff that used to need a real set, a real camera op, and a full shoot day. You can generate every frame of it from a text prompt now, and stitch it into a finished ad before your coffee goes cold. The workflow is almost stupidly simple: → Tell Claude Opus 4.8 the beauty shot you want (dewy skin macro, gloss-on-lips contact, ripple transition, the works) → Claude turns it into a shot-by-shot storyboard plus a prompt for every frame → GPT Image 2 generates the photoreal stills, frame by frame → Seedance 2.0 animates each one into a clip with that buttery slow-mo glide → You drop the clips into HeyOz and assemble the full ad in one place The real unlock is volume. This isn't one hero video. Once the workflow is dialed, you spin up hundreds of variations. Different shades, different models, different hooks, different transitions. The exact creative volume Meta rewards, minus the production cost that used to make it impossible. Old way: one shoot, one look, $10k+, weeks of waiting. New way: a hundred angles, any look, a few dollars each, same afternoon. I wrote up the entire workflow. The Claude storyboard prompt, the GPT Image 2 frame prompts, the Seedance motion settings, the full assembly flow. Completely free, no email gate. Want it? Comment "GLOSS" and I'll send it straight over. (make sure you're following so it can actually reach you)

Ahad Shams

11,067 просмотров • 1 месяц назад

The future of housework just leaked on GitHub and nobody is talking about it. knox byte just open sourced a framework that coordinates swarms of Unitree G1 humanoid robots to clean your entire house on their own. It's called ARGOS. You tell it "clean the bedroom" in plain English and 2+ G1 robots split the room into zones, sweep in parallel, and sync up for the tasks that need four hands like making the bed or moving furniture. The Claude API decomposes your sentence into a task graph. An auction system makes every robot bid on every task based on distance, battery, and current load. The cheapest robot wins. Cooperative jobs go to the cheapest team. Here's what makes this different from every demo video Boston Dynamics keeps teasing: → 12 cleaning tasks baked in sweeping, mopping, wiping, vacuuming, taking out trash, making the bed, changing sheets, moving furniture, sorting items → 3 policy architectures running underneath OpenVLA-7B for language tasks, Diffusion Policy for floor coverage, ACT for dexterous bimanual work → Train it on your own footage record yourself cleaning, run one command, it extracts poses, builds a LeRobot dataset, and LoRA fine-tunes the policy → PEFA protocol for cooperative work Propose, Execute, Feedback, Adjust. If one robot fails halfway through making the bed, the team replans and retries → Full MuJoCo simulation so you test policies before pushing them to real hardware → Silver and cyan terminal dashboard that shows live fleet status, zone maps, task queues, and battery levels in real time The G1 robots talk to each other over CycloneDDS mesh using Unitree's native SDK. No cloud. No middleware. The whole thing runs on a Jetson Orin inside each robot. The wildest part is the training pipeline. Drop cleaning videos into a folder, run argos train ingest, and the framework does the entire pipeline frame extraction, pose estimation, action labeling, HDF5 dataset, fine-tune, evaluate in sim, deploy to robot. One command per stage. Unitree G1s already exist. The framework to make them clean your house just hit GitHub. 52 stars. MIT License. 100% Opensource.

Guri Singh

27,404 просмотров • 2 месяцев назад

Elon Musk gave the entire entertainment industry its expiration date, and he is the one building the thing that kills it. Musk: “My guess is that we see the first compelling half hour, pure AI show next year.” Next year. A complete show generated entirely by AI. No writers. No actors. No cameras. No sets. No crew. No studio. Just a prompt and enough compute to render a reality that never physically existed. And shows are the easy part. Musk: “I say probably we’re maybe three years away from AI does the whole video game.” A show plays the same way every time. A game has to generate a living world that reacts to every decision in real time across every single frame. That is a fundamentally harder class of problem. And Musk put three years on it. Right now a single AAA title takes seven years and half a billion dollars across thousands of engineers and artists just to ship it. Musk is describing a world where one person types a paragraph and gets something comparable. The entire value proposition of a multi-billion dollar industry lives inside that gap. And it closes in thirty-six months. But the prediction is not the story. The person making it is. This is not an analyst speculating from the sidelines. This is the man building the largest AI compute clusters on the planet. The man who built xAI from zero in under two years. The man stacking hundreds of thousands of GPUs into facilities designed to do exactly what he is describing. When Musk says three years, he is not guessing about what someone else might eventually ship. He is reading you a delivery date off his own roadmap. Every media company on Earth is valued on a single assumption. That quality content is expensive and difficult to produce at scale. That one assumption is the structural foundation underneath every studio, every network, and every publisher in existence. Musk is dismantling it with raw compute. The studios still parading thousand-person production teams are not demonstrating strength. They are advertising the exact cost structure that one person with a prompt and a GPU allocation is about to make irrelevant. And it does not stop at entertainment. If AI can generate an interactive world that responds to human input in real time, it can generate anything. Advertising. Architecture. Training simulations. Product design. Every industry built on humans manually constructing visual experiences frame by frame is sitting on the same countdown Musk just read out loud. Now zoom out. Because this is not just an industry story. For the entire history of human civilization, the distance between imagining a world and actually creating one required thousands of people, millions of hours, and billions of dollars. That distance built Hollywood. That distance built the gaming industry. That distance made content scarce and studios powerful. Musk is collapsing that distance to zero. When the gap between imagining something and it existing disappears, every business model built on the difficulty of creation disappears with it. That is not disruption. That is a full inversion of how human beings create. Musk did not make a casual prediction on that podcast. He told you what he is building. He told you the timeline. And he told you which industries do not survive it. The entertainment industry is still debating whether this future is real. Musk is not part of that debate. He is building. And he just told you the delivery date.

Dustin

21,842 просмотров • 14 дней назад

Okay everybody. I want you all to watch this video closely, so that I can show you what utter made up bullshit propaganda this stupid criminal kid keeps posting. It’s all made up garbage and it’s so embarrassingly fake that it’s hilarious. I’ve attached the video here again for ease of reference, plus will timestamp it with screenshots so you can see it for yourselves. This post alleges that this is a video showing IDF snipers shooting a a boy. 1. Firstly, the videos are split, different coloring, rendering, and motions. These two videos are completely unrelated and have been put one on top of the other as simple propaganda. 2. Look at both sides of the staircase. One either side are clothing stores with mannequins. Those are not even real people on the sides. It’s a small group of people coming down then up the stairs but a few mannequins being knocked over from the left of the screen to the center right. You can see the mannequins falling over like cut timber (still image pointing them out attached). One person falls down but isn’t shot as he gets up and walks away unharmed. 3. At 00:09 in the video, the man bottom right takes a few steps up and leans down… to pick up a freaking mannequin. I’ve taken a still to point him out. 4. For the next two seconds he lifts the mannequin by the head and you can clearly see it’s a mannequin. I’ve taken a still to circle the doll. 5. Slow it down and you will see one shot is fired in the unrelated video below, yet 3-4 “people” falling down from the magic bullet that just plays pinball. One real person lays down but gets up. All the others are knocked over mannequins. Hope you all enjoyed the show, and hopefully you laughed as much as I did. This entire thing is just fake bullshit. It’s the only thing these mentally ill and retarded propagandists ever post. GAZAWOOD - the PALLYWOOD saga feel free to do your thing. H/T Stealth Medical

Cheryl E 🇮🇱🇮🇱🇮🇱🎗️

19,203 просмотров • 1 год назад

Google dropped a new AI paper called LUMIERE. It's remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:

Bilawal Sidhu

44,822 просмотров • 2 лет назад

50% cheaper Claude inference with just one line of code change! - Remove → model="claude-opus-4-8" - Add → model="ship-like/claude-opus-4-8" I verified the cost saving in my own terminal by invoking the same Anthropic model with the same prompt. The underlying engineering by Ship is actually interesting, and the patterns can be used in any production LLM stack. Essentially, a trained model is a frozen artifact. Every request performs the same forward-pass, whether it extracts a date or refactors a module, because the compute decision was made at training time, before the request existed. Ship makes that decision at inference time instead. After seeing a request, it searches over executions, involving single models, cascades, ensembles, or harnesses with tools, and serves the cheapest one that will match the reference model's quality. This is not a basic router, because picking a cheaper model per query doesn't ensure the cheaper model preserves the original's behavior, like output shape, tool-call patterns, and refusals. Ship measures this equivalence directly. Outputs stay distributionally indistinguishable from the reference model, not token-identical, since two calls to the same model already differ, but they are indistinguishable in capability and behavior. Of course, some requests execute cheaply and some cost Ship more than the customer pays, but the price per request is still a flat 50% off either way, so the execution-cost variance moves off the application's bill entirely. The video below depicts the cost savings and output in my real invocation, and I partnered with the team to put this together.

Akshay 🚀

63,647 просмотров • 7 дней назад

Contact sheet prompting is the hottest AI video technique right now 🤯 If you've seen this technique blowing up, here's why it works: You feed AI one image, and it generates a grid of consistent shots—same face, same outfit, different angles and poses. Instant storyboarding. Full creative control. No reshoots. But doing it manually is brutal: → Write the prompt from scratch → Generate the contact sheet → Crop each frame by hand → Feed frames into a video model one at a time → Repeat for every single product That's hours of work per campaign. This n8n automation handles everything: → Upload a character image + product image → AI analyzes both and writes the contact sheet prompt → Nano Banana Pro generates a 6-frame grid → System extracts each frame automatically → Kling 2.5 creates smooth transitions between frames → 5 video clips land in Airtable ready to use No manual cropping. No frame-by-frame prompting. No tedious busywork. What you get in Airtable: - AI-generated creative prompt - Hero image (model + product) - Full 6-frame contact sheet - 5 cinematic video clips - Approval gates before each step All inside n8n + Airtable. Contact sheet prompting on complete autopilot. I recorded a 20-minute Loom showing exactly how I built this. Want the walkthrough + the full n8n workflow + Airtable base? > Like this post > Comment "CONTACT" And I'll send it over (must be following so I can DM)

Mike Futia

24,575 просмотров • 6 месяцев назад

I just built a skill that lets Claude Code watch & analyze ANY video 🤯 Drop in any video file — UGC ads, competitor Meta ads, organic TikToks, screen recordings — and Claude hands you back a full creative teardown. All inside Claude Code. Perfect for media buyers and creative strategists who reverse-engineer competitor ads every week — and lose half a day doing it by hand. If your creative process starts with studying what's already working, you're scrubbing through competitor ads frame by frame, pausing to write down every hook, screenshotting the on-screen text, and by the tenth video you can't remember what made the first one land... This skill solves it: → Drop any video file into Claude Code → Skill routes it through the Gemini API for native video understanding → Returns a full creative teardown — hook breakdown, target audience, angle, beat-by-beat, on-screen text verbatim → Surfaces the steal-worthy patterns you can apply to your own creative → Same skill works on UGC ads, produced video ads, organic TikToks, and Loom recordings No manual scrubbing. No pausing every 5 seconds. No $200/mo ad intelligence platform. What you get: → Native video understanding via Gemini (not just transcripts) → Structured analysis — hook, angle, audience, pain point, CTA → Verbatim on-screen text and dialogue with timestamps → Hook variations generated directly from competitor ads → About 27 cents per 30-minute video Built 100% in Claude Code with the Gemini API. I recorded a full breakdown showing exactly how I built this, and I'm giving away the skill for free. Want the skill? > Like this post > Comment "CLAUDE" And I'll send it over (must be following so I can DM)

Mike Futia

41,284 просмотров • 1 месяц назад

I just built a $10K/month creative strategist inside Claude Code 🤯 Give it your competitor Facebook page URLs → it scrapes their ads, watches every video with AI, and delivers a data-backed creative brief with 10 ad concepts in your brand voice. All inside Claude Code. Perfect for DTC brands and agencies who are still manually scrolling the Meta Ad Library, screenshotting ads into Google Docs, and guessing at what's working. If you're spending hours every week pulling competitor ads one by one, watching videos to figure out the hook, copying notes into a brief, and rewriting concepts from scratch every time... This system eliminates the entire loop: → Apify scrapes your competitors' active ads from Meta Ad Library (video + image) → Downloads every creative asset locally → Gemini watches each video and analyzes the hook, angle, visual format, copy framework, CTA, and emotional trigger → Runs the full batch and finds the patterns that repeat across 3+ ads → Claude generates 10 ad concepts using the proven mechanics, matched to your brand voice No manually scrolling the Ad Library. No screenshotting ads into docs. No guessing which hooks are actually working. What you get: → Individual creative breakdowns for every competitor ad (7 dimensions each) → A pattern report showing which hooks, formats, and triggers keep repeating → 10 ready-to-brief ad concepts traced back to real competitor data → A reusable system — new competitors, new brief, same pipeline The research that takes your team a full day now runs in 15 minutes for ~$3 in API costs. Built 100% in Claude Code with Apify + Gemini. I put together a full playbook showing you can build the entire thing step-by-step from scratch. Want the playbook for free? > Like this post > Comment "ADS" And I'll send it over (must be following so I can DM)

Mike Futia

54,667 просмотров • 4 месяцев назад