Introducing Moto: Latent Motion Token as the Bridging Language... for Robot Manipulation. Motion prior learned from videos can be seamlessly transferred to robot manipulation. Code and model released! Yi CHEN Yuying Ge Yixiao Ge Mingyu Ding Ying Shanshow more

Xihui Liu
15,030 görüntüleme • 1 yıl önce
✨ Any static 3D assets ➡️ 4D dynamic worlds.... Introducing CHORD, a universal framework for generating scene-level 4D dynamic motion from any static 3D inputs. It generalizes surprisingly well across a wide range of objects 🤯 and can even be used to learn robotics manipulation policy 🤖! Project page: Dive deeper in a 🧵: 1/nshow more

Chen Geng
43,469 görüntüleme • 7 ay önce
30 minutes of video. Robot learns the task. Open-source,... end-to-end. An open-source framework for training robot policies from only 30 minutes of human egocentric videos captured via Meta Aria glasses: Achieving zero-shot transfer to robots without any robot data collection. The method relies on Interaction-Centric Tokens that encode hand-object spatial relationships invariant to embodiment and viewpoint, supplemented by auxiliary objectives like object motion prediction and latent consistency to extract richer supervision signals from the same data. HumanEgo demonstrates strong cross-embodiment, cross-environment performance on bimanual tasks, outperforming baselines like ACT and teleop data while being trainable on a single RTX 4090 GPU. Thanks for sharing, Zhi (Leo) Wang. 📌 Website: Paper: Code: Video: ——- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
17,077 görüntüleme • 3 ay önce
Building robots that can effectively operate alongside human workers... is difficult. 🛠️ Advances in open-source physics, open foundation models, and frameworks are helping accelerate physical #AI deployment. ✔️ Newton Physics Engine, an open-source GPU-powered simulation built on OpenUSD, speeds up robot learning for advanced manipulation and mobility. ✔️ NVIDIA Cosmos Reason, an open reasoning vision language model, gives robots the ability to think like humans using prior knowledge, common sense and physics ✔️NVIDIA Isaac GR00T N1.6, an open robot foundation model, enables humanoids to understand ambiguous instructions Leading robotics developers including Agility Robotics, Lightwheel, Mentee Robotics, UniversalRobots, and Wandelbots are adopting simulation technologies and libraries to accelerate physical AI development and deployment. Omniverse Ambassador Dylan Tobin built an AI chatbot trained on Isaac Sim workflows, helping devs navigate Omniverse faster. Read the full blog 👉show more

NVIDIA
48,500 görüntüleme • 11 ay önce
ZACHXBT FLAGS POSSIBLE INSIDER ACTIVITY ON $RAVE On-chain investigator... ZachXBT has raised concerns about alleged insider manipulation of the RaveDAO token, posting that insiders appear to control more than 90% of $RAVE's supply. In his post, ZachXBT pointed to what he described as pump-and-dump activity originating from Binance, @bitgetglobal, and Gate.io, and called on Binance's He Yi and Bitget CEO Gracy Chen @Bitget to investigate and offboard those involved. He also offered a $10,000 bounty for whistleblowers who come forward with evidence related to the alleged scheme. Bitget CEO Gracy Chen publicly responded shortly after, confirming the exchange had opened an investigation into $RAVE. $RAVE is up over 10,000% in the past 30 days, with the token rallying another 44% on Saturday.show more

BSCN
37,739 görüntüleme • 4 ay önce
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
12,379 görüntüleme • 2 ay önce
TESLA HALTED MODEL S AND MODEL X PRODUCTION TO... BUILD AN ARMY OF OPTIMUS ROBOTS The Fremont assembly line was torn down in 46 days. In its place, Tesla is building a line for humanoid production, aiming for a million units a year A humanoid robot is a body shaped like a human. Physical AI is the intelligence that controls that body Walking and making coffee is often just imitation learning from a scripted routine. But once the environment shifts, the learned trick stops working Language models had the entire internet to train on. Robotics has nothing close to that scale of data, which is why one giant brain hasn't worked for anyone yet The industry is moving toward modularity instead - separate models for vision, movement, and planning, each improved on its own The real question is no longer whether a robot can move impressively. It's whether it can pull its sensors into one picture of the world and adapt to whatever wasn't scripted for itshow more

iamigorekk
22,205 görüntüleme • 16 gün önce
𝗗𝗼𝗻'𝘁 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗲 𝗿𝗼𝗯𝗼𝘁 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹𝘀. 𝗦𝘁𝗲𝗲𝗿 𝘁𝗵𝗲𝗺 𝘄𝗶𝘁𝗵 𝗵𝘂𝗺𝗮𝗻... 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘀𝘁𝗲𝗮𝗱, 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗯𝗮𝘀𝗲 𝗽𝗼𝗹𝗶𝗰𝘆 Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging. A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe. In our new paper, 𝗙𝗹𝗼𝘄𝗗𝗔𝗴𝗴𝗲𝗿, we take a different approach: 𝗜𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹, 𝘄𝗲 𝗹𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗲𝗲𝗿 𝗶𝘁 𝗳𝗿𝗼𝗺 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀. The key idea is 𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘃𝗲𝗿𝘀𝗶𝗼𝗻: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities. Across simulation and real robots, FlowDAgger: 📈 Learns from only 5–20 human intervention episodes 🏆 Outperforms supervised fine-tuning and latent-space reinforcement learning 🤖 Works across VLAs, diffusion policies, and world-action models ✔️ Provides reliable improvements without modifying the pretrained policy We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections. 📄 Paper: 🌐 Project: 💻 Code: This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and Andrey Kolobov at Microsoft Research and Maya Cakmak at University of Washingtonshow more

Oier Mees
13,032 görüntüleme • 1 ay önce
Introducing VL-JEPA: Vision-Language Joint Embedding Predictive Architecture for streaming,... live action recognition, retrieval, VQA, and classification tasks with better performance and higher efficiency than large VLMs. • VL-JEPA is the first non-generative model that can perform general-domain vision-language tasks in real-time, built on a joint embedding predictive architecture. • We demonstrate in controlled experiments that VL-JEPA, trained with latent space embedding prediction, outperforms VLMs that rely on data space token prediction. • We show that VL-JEPA delivers significant efficiency gains over VLMs for online video streaming applications, thanks to its non-autoregressive design and native support for selective decoding. • We highlight that our VL-JEPA model, with an unified model architecture, can effectively handle a wide range of classification, retrieval, and VQA tasks at the same time. by Delong Chen (陈德龙) Mustafa Shukor Théo Moutakanni Willy Jade Lei Yu Tejaswi Kasarla Allen Bolourchi Yann LeCun Pascale Fungshow more

Pascale Fung
90,144 görüntüleme • 8 ay önce
🦿Xpeng showed a humanoid robot called IRON whose movement... looked so human that the team literally cut it open on stage to prove it is a machine. IRON uses a bionic body with a flexible spine, synthetic muscles, and soft skin so joints and torso can twist smoothly like a person. The system has 82 degrees of freedom in total with 22 in each hand for fine finger control. Compute runs on 3 custom AI chips rated at 2,250 TOPS (Tera Operations Per Second), which is far above typical laptop neural accelerators, so it can handle vision and motion planning on the robot. The AI stack focuses on turning camera input directly into body movement without routing through text, which reduces lag and makes the gait look natural. Xpeng staged the cut-open demo at AI Day in Guangzhou this week, addressing rumors that a performer was inside by exposing internal actuators, wiring, and cooling. Company materials also mention a large physical-world model and a multi-brain control setup for dialogue, perception, and locomotion, hinting at a path from stage demos to service work. Production is targeted for 2026, so near-term tasks will be limited, but the hardware shows a serious step toward human-scale manipulation.show more

Rohan Paul
3,802,543 görüntüleme • 9 ay önce
[Most robots react. This one thinks a step ahead.]... Ant Group's Robbyant just published LingBot-VA 2.0 — a video-action foundation model built from scratch for robot control, not fine-tuned from a video generator. The usual approach takes a video generator made for content creation and bolts a robot policy onto it. LingBot-VA 2.0 argues that's the wrong starting point, and pretrains the whole causal stack natively instead. What stands out: → Foresight Reasoning — the robot predicts the next action chunk while executing the current one, then overwrites the imagined frame with the real observation. Prediction and execution stop waiting on each other. → 927 ms → 142 ms per chunk, across four cumulative optimizations. That lifts asynchronous control from 35 Hz to 225 Hz — a 6.5× speedup. → One shared latent space. A semantic visual-action tokenizer puts world states and actions in the same coordinates, so unlabeled web video carries action-relevant signal. → Sparse MoE video stream — 128 experts, top-8 routing. Roughly 2.5B of ~15.3B parameters fire per token. → Few-shot by design — adapts from 10–15 demonstrations, and a human demo video can replace the text instruction entirely. Full breakdown: Paper: Project Page: Robbyant Ant Groupshow more

Marktechpost AI
196,499 görüntüleme • 1 ay önce
The Hidden Language of Diffusion Models paper page: tackle... the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a pseudo-token that is a sparse weighted combination of tokens from the model's vocabulary, with the objective of reconstructing the images generated for the given concept. Applied over the state-of-the-art Stable Diffusion model, this decomposition reveals non-trivial and surprising structures in the representations of concepts. For example, we find that some concepts such as "a president" or "a composer" are dominated by specific instances (e.g., "Obama", "Biden") and their interpolations. Other concepts, such as "happiness" combine associated terms that can be concrete ("family", "laughter") or abstract ("friendship", "emotion"). In addition to peering into the inner workings of Stable Diffusion, our method also enables applications such as single-image decomposition to tokens, bias detection and mitigation, and semantic image manipulationshow more

AK
41,830 görüntüleme • 3 yıl önce
Kling 2.6 Motion Control is absolutely insane 🤯 Take... any reference video and transfer the exact motion onto an AI character: full-body sync, facial expressions, hand gestures, everything. All with just a few clicks. Perfect for e-comm brands and agencies creating AI video ads that don't look like AI. Here's the problem: AI-generated video ads still look robotic. The movements are stiff, the expressions are flat. Your audience clocks it as AI instantly and keeps scrolling. Kling 2.6 Motion Control fixes it: → Start with any reference clip (stock footage, existing UGC, motion reference) → Upload to Kling → Map the exact movement onto any AI character → Full-body motion, hand gestures, facial expressions—all transferred → Generate up to 30 seconds of video No stiff AI movements, no uncanny valley, no instant "skip this ad" reaction. What this unlocks: - Use one winning UGC motion → swap in different AI creators - Pull reference clips from anywhere → generate branded variations - Create dynamic AI video ads with real human movement - Test multiple "creators" without filming anyone new I recorded a quick walkthrough showing how to do this step-by-step. Want access? > Comment "KLING" > Like this post And I'll send it over (must be following so I can DM)show more

Mike Futia
25,247 görüntüleme • 7 ay önce
You pretrained a robot policy on millions of camera... frames. Now you want to add a force sensor. Do you really have to retrain on everything from scratch? MuSe adds a force-torque sensor to a frozen vision-only policy using a tiny amount of contact data. - lifts contact-rich task success (peg insertion 60% to 87%, vase wiping 33% to 77%) - fuse the new sensor both early (shared token space) and late (cross-attention), train the policy as a world model that predicts future video, future force, and actions together, and replay old vision-only data with the force input masked to prevent forgetting. - adding new sensor modalities to pre-trained World Model can be cheap & improve performance significantlyshow more

Vai Viswanathan
36,059 görüntüleme • 2 ay önce
🛠️ What if a robot could invent its own... tools. And teach itself how to use them? That’s exactly what VLMgineer does: a new framework that lets Vision Language Models (VLMs) design physical tools and the actions to use them, entirely on their own. No templates. No human demonstrations. Just raw, AI-driven creativity. Why it matters ✅ Co-designs tools and actions together using VLMs, ensuring tight coupling between form and function ✅ Uses VLM-guided evolution (not random search) to refine designs intelligently ✅ Outperforms human-designed tools by +64.7% in task success across 12 RoboToolBench challenges ✅ Produces better-than-everyday tools for real manipulation tasks—measured in success rate and elegance It builds on the emerging trend of large-model-guided evolutionary design (like Eureka and AlphaEvolve) and brings it into physical robotics. It opens the door to general-purpose, automated hardware design, no strong priors needed. Code & paper: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
13,984 görüntüleme • 8 ay önce
Today may be the ImageNet moment for robotics. RT-X:... the largest open-source robot dataset ever compiled, across 33 institutes, 22 robot hardware, 527 skills, and 1M episodes. Why is robotics lagging so far behind NLP, vision, and other AI domains? Data scarcity is the main culprit to blame, among other difficulties. Unlike text, images, and videos, you cannot download mass amounts of onboard robot control data from the internet. They simply don't exist in the wild. 11 yrs ago, ImageNet kicked off the deep learning revolution. 3-4 yrs ago, internet-scale data fueled the first GPTs and Diffusions that define this era of foundation models. I think 2023 is finally the year for robotics to scale up. Robot foundation models like VIMA ( my team's work at NVIDIA) and RT-1/2 ( Google DeepMind's effort) are extremely data hungry. While massively parallel simulations like NVIDIA IsaacGym & Omniverse can alleviate the problem to some extent, it's still not quite enough to bridge the gap to the messy, physical world. This new dataset is not just a technical contribution. I also see it as a commendable effort to overcome institutional bureaucracies and unite researchers from around the world to tackle a grand challenge together. Robotics will be the final holy grail that we capture in AI. We are not there yet, but ascending in the right gradient direction. RT-X website: Launch blog:show more

Jim Fan
265,061 görüntüleme • 2 yıl önce
Robot Utility Models (RUMs) enable basic tasks – door... opening, drawer opening, object reorientation, etc. – at ~90% accuracy without ANY finetuning (i.e. zero-shot) in unseen new environments. Fully open source!!! models, data, code & hw. We think this is super exciting, why?👇 1. Unlocks many practical home utility tasks that often involve these basic tasks as part of an action chain. “Go get me a fork” involves opening the kitchen door and then opening the cutlery drawer. 2. This works well **zero-shot in unseen and new** environments, which is practically a huge deal. Turn the robot on, and get going. 3. The recipe for building a new model is fairly generic, and we think with a bit more refinement this can be a general recipe to build many more Utility models. More details and access 👇show more

Mahi Shafiullah 🏠🤖
89,535 görüntüleme • 2 yıl önce
South Korea's local governments are deploying around 7,000 AI-robot... dolls to seniors and dementia patients. The $1,800 robot doll by Hyodal can hold full conversations to tackle loneliness and remind users to take medication. Dystopian, yes, but the data is fascinating: 1. Studies (with over 9,000 users) found that depression levels reduced from 5.73 to 3.14, and medicine intake improved from 2.69 to 2.87. 2. The doll comes with a companion app and web monitoring platform for caretakers to monitor remotely. 3. Safety features are installed to alert when no movement has been detected for a certain period, essentially always watching the user. 4. The doll also offers touch interaction, 24-hour voice reminders, check-ins, voice messages, a health coach, quizzes, exercise, music, and more. 5. Caregivers have access to the app, allowing them to send/receive voice messages, make group announcements, and monitor motion detection. I'd definitely have some privacy and data collection concerns here before handing this off to my family, but the product actually seems really cool. Will be interesting to watch the data to see if this idea has legs. Keep in mind, SK has a rapidly aging population and one of the world's lowest birth rates, so it makes sense for the local governments to be early adopters here.show more

Rowan Cheung
420,862 görüntüleme • 2 yıl önce
𝗣𝗼𝗽𝘂𝗹𝗮𝗿 𝗼𝗽𝗶𝗻𝗶𝗼𝗻: "𝗝𝘂𝘀𝘁 𝗴𝗲𝗻𝗲𝗿𝗮𝘁𝗲 𝗺𝗼𝗿𝗲 𝘀𝗶𝗺𝘂𝗹𝗮𝘁𝗶𝗼𝗻 𝗱𝗮𝘁𝗮." After working... with many 𝗿𝗼𝗯𝗼𝘁 𝗺𝗮𝗻𝗶𝗽𝘂𝗹𝗮𝘁𝗶𝗼𝗻 teams who've fallen into the simulation trap, here's what I've learned: Simulation teaches your robot to be really, really good at simulation. Unlike blind locomotion policies that can get away with sim-to-real transfer because they rely mainly on proprioception and contact forces, 𝘃𝗶𝘀𝗶𝗼𝗻-𝗴𝘂𝗶𝗱𝗲𝗱 𝗺𝗮𝗻𝗶𝗽𝘂𝗹𝗮𝘁𝗶𝗼𝗻 𝗶𝘀 𝗲𝘅𝘁𝗿𝗲𝗺𝗲𝗹𝘆 𝘀𝗲𝗻𝘀𝗶𝘁𝗶𝘃𝗲 𝘁𝗼 𝘃𝗶𝘀𝘂𝗮𝗹 𝗱𝗼𝗺𝗮𝗶𝗻 𝗴𝗮𝗽. The subtle differences accumulate: - Simulated friction vs real surface textures - Perfect lighting vs shadows, reflections, glare - Ideal object geometries vs manufacturing tolerances - Instantaneous sensor readings vs real-world noise and latency - Clean backgrounds vs cluttered, dynamic environments 𝗧𝗵𝗲 𝗰𝗹𝗮𝘀𝘀𝗶𝗰 𝗽𝗿𝗼𝗴𝗿𝗲𝘀𝘀𝗶𝗼𝗻: Week 1: "Our model works perfectly in sim!" Week 2: "Let's collect some real data to fine-tune." Week 3: "The real data completely contradicts what the sim taught..." Week 4: "Okay, let's collect way more real data." Month 2: "We basically need to retrain from scratch." 𝗧𝗵𝗲 𝗽𝗮𝗶𝗻𝗳𝘂𝗹 𝘁𝗿𝘂𝘁𝗵: There's no shortcut to real-world data collection for vision-based manipulation. Simulation is amazing for debugging, prototyping, safety testing, and of course to supplement your real data. But it's not a substitute for understanding how your robot actually behaves in the actual environment. 𝗪𝗵𝗮𝘁 𝘄𝗼𝗿𝗸𝘀: Use simulation strategically - for exploring edge cases, testing safety boundaries, and rapid iteration. But build your production models on real data from real environments. The teams that succeed treat simulation as a powerful tool, not a magic solution. This is why Neuracore focuses on making real-world data collection so much easier and faster. Because the physics of your actual environment can't be simulated away. 𝗪𝗼𝗿𝗹𝗱 𝗺𝗼𝗱𝗲𝗹𝘀, 𝘆𝗼𝘂 𝘀𝗮𝘆? 𝗪𝗲𝗹𝗹, 𝗽𝗲𝗿𝗵𝗮𝗽𝘀 𝗺𝗼𝗿𝗲 𝗼𝗻 𝘁𝗵𝗮𝘁 𝗶𝗻 𝗮𝗻𝗼𝘁𝗵𝗲𝗿 𝗽𝗼𝘀𝘁! 𝗪𝗵𝗮𝘁'𝘀 𝗯𝗲𝗲𝗻 𝘆𝗼𝘂𝗿 𝗲𝘅𝗽𝗲𝗿𝗶𝗲𝗻𝗰𝗲 𝘄𝗶𝘁𝗵 𝘀𝗶𝗺-𝘁𝗼-𝗿𝗲𝗮𝗹 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿? 𝗛𝗮𝘀 𝗶𝘁 𝘄𝗼𝗿𝗸𝗲𝗱 𝗮𝘀 𝘄𝗲𝗹𝗹 𝗮𝘀 𝗲𝘅𝗽𝗲𝗰𝘁𝗲𝗱?show more

Stephen James
31,009 görüntüleme • 1 yıl önce
We have released Seedance 2.0. Due to the 2500-character... limit, please translate the following prompts into Chinese before use. [Technical Specs] Generate a 10-second, 16:9, 720p cinematic video. Smooth continuous camera motion with no cuts. The overall pacing is fast and tightly compressed, with rapid escalation from start to finish. Audio evolves quickly from a high-performance engine idle into intricate mechanical shifting and clicks, culminating in a soft electronic chime and the distinct sound of a "mwah" blowing kiss. [Global Constraints] Only the evolving mechanical character appears; no other humans or characters. All transformations must follow physical logic and maintain structural continuity. No object should pass through or intersect with other solid objects. Every robotic component must originate from visible parts of the Porsche 911 (doors, hood, wheels, chassis) through unfolding, splitting, or reconfiguration. [Scene Setup — 0:00–0:01] A sleek, metallic silver Porsche 911 sits on a rain-slicked futuristic city street at night, neon lights reflecting off its polished surface. The camera starts at a low-angle front-quarter view and begins a fast, smooth tracking-arc towards the side. [Rapid Transformation Initiation — 0:01–0:03] Transformation triggers instantly. The car’s suspension drops, and the frame begins to fracture into a complex grid of panels. The doors swing open and begin to segment into articulated arm structures. The front hood splits down the center, folding inward to reveal a glowing internal core. The headlights flicker and start to reorient as the "eyes." [Accelerated Feminine Reconfiguration — 0:03–0:07] The mechanical action is dense, overlapping, and fluid, emphasizing graceful but powerful motion. Lower Body: The rear wheels and wheel arches split and rotate downward, reassembling into slender, high-heeled mechanical legs. Torso: The roof and rear engine cover slide and compress, forming a sleek, curvaceous hourglass torso that retains the car’s aerodynamic lines. Arms & Hands: The side mirrors and door panels unfold into delicate but strong hands and fingers. Head: The front bumper and emblem area segment and rise, folding into a feminine-shaped head with a sleek metallic "helmet" visor. [Logical Transformation Constraints — No Spontaneous Appearance] The robot’s "skin" is composed of the car's outer silver panels. The internal frame and wiring emerge from the engine and undercarriage. No parts appear out of thin air; every joint is a reconfigured automotive component. [Transformation Completion — 0:07–0:08.5] The robot stands tall and elegant. The silver panels lock into place with a satisfying "click," revealing glowing blue LED accents in the seams. The silhouette is clearly feminine, humanoid, and sophisticated, reflecting the premium design of the original vehicle. [Final Hero Ending — 0:08.5–0:10] As the robot stabilizes, the camera performs a rapid, smooth zoom-in (Dolly-In) directly to her face. The robot tilts its head slightly, and the optic sensors (eyes) brighten. It brings its mechanical hand to its metallic lips and performs a graceful blowing kiss (fly-kiss) gesture toward the camera. The video ends with a close-up of the face, capturing the reflection of neon lights in its visor just as the kiss is released. [Cinematography Notes] Continuous Motion: No cuts or fades; the camera must transition from the car-tracking shot to the face-zoom seamlessly. Material Consistency: The robot must maintain the exact metallic silver paint, texture, and reflections of the Porsche. Energy: The transformation should feel high-energy and "force-driven," while the final gesture is soft and charismatic.show more

underwood
19,462 görüntüleme • 5 ay önce
FABLE 5 + HIGGSFIELD JUST KILLED THE $35,000 WEB... AGENCY. The same animated website that used to cost between $6,000 and $35,000 can now be built in a single session for around $12 in AI credits. Claude Code handles the website. Higgsfield creates the visuals. Together, they can build a complete scroll-driven website from a simple prompt. Claude writes the layout, GSAP ScrollTrigger animations, Lenis smooth scrolling, responsive pages, and checks everything before you ship. Higgsfield generates the hero videos, cinematic transitions, ambient loops, thumbnails, and every visual asset you need. By the end of one session, you have a fully animated website with cinematic motion, smooth scrolling, optimized assets, responsive layouts, and polished visual effects like film grain, particles, vignette, glass cards, and color tints. Getting started only takes a few minutes. Add Higgsfield as an MCP server inside Claude Code, complete the OAuth login once, and Claude can generate and pull videos directly into your project without manually exporting anything. The prompts are simple. Give Claude your project brief and ask it to script the entire scroll experience. Tell it to generate a hero video and supporting clips for every section. Ask it to add the finishing touches like film grain, particles, glass cards, and scroll pacing. Then let it review the site, improve loading speed, fix mobile layouts, and rewrite anything that doesn’t work. This replaces work that usually looks like this: A $6,000-$35,000 web agency. An $800-$2,000 motion designer. A $2,000-$10,000 front-end developer. And weeks of back-and-forth before launch. Now it’s Fable 5, Higgsfield, a subscription, a few dollars in AI credits, and one session. The pipeline used to be the advantage. Now it’s just a prompt. Full breakdown in the article below.show more

MIKE
141,788 görüntüleme • 1 ay önce