Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

How can we leverage diverse human videos to improve robot manipulation? Excited to introduce EgoVLA — a Vision-Language-Action model trained on egocentric human videos by explicitly modeling wrist & hand motion. We build a shared action space between humans and robots, enabling seamless transfer. With some robot demos, EgoVLA...

58,800 Aufrufe • vor 1 Jahr •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more details

Oier Mees

12,379 Aufrufe • vor 2 Monaten

𝗗𝗼𝗻'𝘁 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗲 𝗿𝗼𝗯𝗼𝘁 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹𝘀. 𝗦𝘁𝗲𝗲𝗿 𝘁𝗵𝗲𝗺 𝘄𝗶𝘁𝗵 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘀𝘁𝗲𝗮𝗱, 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗯𝗮𝘀𝗲 𝗽𝗼𝗹𝗶𝗰𝘆 Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging. A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe. In our new paper, 𝗙𝗹𝗼𝘄𝗗𝗔𝗴𝗴𝗲𝗿, we take a different approach: 𝗜𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹, 𝘄𝗲 𝗹𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗲𝗲𝗿 𝗶𝘁 𝗳𝗿𝗼𝗺 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀. The key idea is 𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘃𝗲𝗿𝘀𝗶𝗼𝗻: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities. Across simulation and real robots, FlowDAgger: 📈 Learns from only 5–20 human intervention episodes 🏆 Outperforms supervised fine-tuning and latent-space reinforcement learning 🤖 Works across VLAs, diffusion policies, and world-action models ✔️ Provides reliable improvements without modifying the pretrained policy We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections. 📄 Paper: 🌐 Project: 💻 Code: This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and Andrey Kolobov at Microsoft Research and Maya Cakmak at University of Washington

Oier Mees

13,032 Aufrufe • vor 1 Monat

Imagine you go to a store and you want to buy candy. The shopkeeper knows you're a real kid because they can see you standing right there. Now imagine you send a robot to buy candy for you. The shopkeeper looks at the robot and thinks: wait, who sent this? Is this robot allowed to buy candy? What if someone else's robot pretends to be yours and steals your candy money? That's basically what's happening with AI right now. Companies like Visa let people buy things all over the world. But now, smart computer robots (AI agents) want to buy things too. Shop around, compare prices, even pay for stuff. Visa looked at this and said: nope, not yet. Because they have no way to check if the robot is real, who it belongs to, or if it's allowed to spend that money. The problem is that all the rules we have for checking identity - showing your ID, scanning your face, typing your password - only work for humans. Robots can't do any of that. Worse, bad robots can actually copy and fake human identities really well. So Evin McMullen evin, Billions Network co-founder and CEO, says we need a new kind of ID system. One where you can prove something is true without showing all your private stuff. Like proving you're tall enough for a ride without telling anyone your exact height. That's called zero-knowledge proof. And for the robots specifically, we need something called KYA - Know Your Agent. It's like giving every robot its own ID card that says: this is who I am, this is what I'm allowed to do, and this is the human responsible for me. Until we build that, the robot economy can't really get going. Here is Evin’s Thought Leader article at Silicon Valleys Journal

Billions Network

21,775 Aufrufe • vor 6 Monaten