Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

What if robots could improve themselves by learning from their own failures in the real-world? Introducing 𝗣𝗟𝗗 (𝗣𝗿𝗼𝗯𝗲, 𝗟𝗲𝗮𝗿𝗻, 𝗗𝗶𝘀𝘁𝗶𝗹𝗹) — a recipe that enables Vision-Language-Action (VLA) models to self-improve for high-precision manipulation tasks. PLD couples real-world residual reinforcement learning with standard supervised fine-tuning — letting robots discover, recover,...

186,079 Aufrufe • vor 11 Monaten •via X (Twitter)

35 Kommentare

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

Supervised fine-tuning (SFT) made today’s strongest robot foundation models in applications — but every gain costs fleets of human teleoperators. PLD replaces them with residual RL specialists that only “take over” when the base policy (any VLA architecture) fails. 🧠 Probe the failure → Learn recovery behavior → Distill back into the generalist. And it works: ✅ Minimal human guidance ✅ ~99 % success on LIBERO ✅ +50 % gain in SimplerEnv ✅ 100 % success on real Franka & YAM arms ✅ 1 hr continuous GPU insertion w/o human reset

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

To perform real-world RL, the algorithm must be super sample-efficient. PLD builds on a series of design choices that make off-policy RL actually practical for robots: 1️⃣ Roll out the base policy longer to flood the demo buffer and oversample successful trials — (RLPD, 2️⃣ Warm up the critic with calibrated Q-learning loss — (Cal-QL, 3️⃣ Carefully choose small initial action scales to prevent RL from sabotaging base policy performance during RL's early stage exploration; 4️⃣ Randomize initial states distribution via base-policy rollouts — unlike JSRL, we skip the scheduler to retain recovery capability (JSRL,

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

Why does it generalize? Because PLD collects policy-aligned data — not "random" teleop demos. It learns from the exact states the model visits and fails in. A model that trains on its own distribution forgets less and adapts faster. Try the interactive demos on our website to get more understanding👇

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

Very fortunate to work with this incredible team: @DarthUtopian, @andy_peng05, @HaoruXue, @TairanHe99, @yuqi_xie5, @fengyuan_hu, @jimmyyhwu, @zhengyiluo, @DrJimFan, @GuanyaShi, @yukez 🔗 Check out more demos on our project website: ❤️ We are also deeply grateful for the help and support from @JasonJZLiu, @_tonytao_, @qiyang_li, @letian_fu, @AjayMandlekar, @youliangtan, @Haoyu_Xiong_, @CharlesXu0124, @guanzhi_wang

Profilbild von Abhishek Ramnath
Abhishek Ramnathvor 11 Monaten

Is human in the loop correction also possible along with RL corrections?

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

PLD aims to minimize human guidance. But having human in the loop is definitely possible (e.g., ConRFT,

Profilbild von Ted Xiao
Ted Xiaovor 11 Monaten

Impressive results, nice work!

Profilbild von Chyna
Chynavor 11 Monaten

Love this. Robots getting better by failing in the real world feels like the most human way to learn.

Profilbild von Hein Van Hoof 🇧🇪
Hein Van Hoof 🇧🇪vor 11 Monaten

The RAM needs to go in Slot 1 and 3, not 1 and 2. OK, then proceed to put it in Slot 3 and 4. 😅

Profilbild von tunglinwood
tunglinwoodvor 11 Monaten

Hi, any plan to open source the project code?

Profilbild von Zi-ang Cao
Zi-ang Caovor 11 Monaten

Congrats!!!

Profilbild von Youssef El Manssouri
Youssef El Manssourivor 11 Monaten

This could democratize robotics development. Small teams could deploy basic robots that improve themselves rather than needing massive training infrastructure.

Profilbild von Minghuan Liu
Minghuan Liuvor 11 Monaten

Congrats! What are the key ingredients you observed during real-world RL?

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

Thanks Minghuan! From my experience, the most critical ingredient for real-world RL is the prior — i.e., the base policy. My take is: If human budget is limited, spend nearly all of it on collect teleop data to pretrain a strong base model. You can see this on my website — there’s a panel visualizing the residual policy output. When the base is good, the residual outputs near-zero most of the time and only corrects at key steps. Early on, I tried a trivial ResNet+MLP base trained with a simple BC loss — RL completely failed, even on basic goal reaching. After switching to ACT and π₀, RL learning became dramatically faster and more stable.

Profilbild von Minghuan Liu
Minghuan Liuvor 11 Monaten

Awesome! Further question: what kind of ability do you expect the base model to cover (for example, which kind of generality)? And what kind of ability does RL deeply improve? Also, is real-world reward design as tough as the ones in training legged robots?

Profilbild von satvik
satvikvor 11 Monaten

Really cool

Profilbild von Himanshu Kumar
Himanshu Kumarvor 11 Monaten

That's a fascinating concept, Wenli! Imagine robots actually learning from their mistakes – it's like a whole new level of "jugad", isn't it?

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

Thanks Himanshu! Exactly, it is all about enabling robots to learn and adapt from their own failures, turning “jugaad” into a systematic self-improvement loop.

Profilbild von Sina G
Sina Gvor 11 Monaten

Hey I would love to use this for my robotics company let’s connect!

Profilbild von Heng (Alfredo)Zhang
Heng (Alfredo)Zhangvor 11 Monaten

Amazing work!

Profilbild von Astrid Wilde 🌞
Astrid Wilde 🌞vor 11 Monaten

@apagajewski of interest

Profilbild von 쇼팽.
쇼팽.vor 11 Monaten

This is a very impressive work, but you cant say that you achieved 99% on Libero when you didnt include the numbers for libero long and 90..

Profilbild von MR
MRvor 6 Monaten

Hello, do you plan to release the code?

Profilbild von Tommy AutoTech
Tommy AutoTechvor 11 Monaten

PLD's 'learn from failure' approach is exactly what warehouse automation needs. Watching robots refine skills in real-time? That's how you build precision. #Automation #Robotics

Profilbild von Clemens Marschner
Clemens Marschnervor 11 Monaten

Good stuff!

Profilbild von Binh
Binhvor 11 Monaten

@nilscmr

Profilbild von Zheyuan Hu
Zheyuan Huvor 11 Monaten

Congrats! Amazing work for real world RL + VLA

Profilbild von Eddy Xu
Eddy Xuvor 11 Monaten

very nice

Profilbild von David Held
David Heldvor 11 Monaten

Where can I find the paper?

Profilbild von Wenli Xiao
Wenli Xiaovor 11 Monaten

The arXiv version will probably come out next week

Profilbild von Yilun Chen
Yilun Chenvor 11 Monaten

Great work, congrats!

Profilbild von wang albert
wang albertvor 4 Monaten

Great work!Would you release the code recently?

Profilbild von Chuck Petras
Chuck Petrasvor 11 Monaten

@BrianRoemmele

Profilbild von Jermaine freejack
Jermaine freejackvor 11 Monaten

This is fascinating. I’m exploring metaphor-driven wellness tech and emotional UX—curious how PLD handles edge cases where failure is emotionally or physically ambiguous. Can VLA models learn from subtle, non-binary outcomes like hesitation or partial success?

Profilbild von Abdi Negese
Abdi Negesevor 5 Monaten

is the project open-sourced? would love to replicate it

Ähnliche Videos

🚨 BREAKING: Microsoft's first robotics foundation model! 🤯 Microsoft just announced Rho-alpha (ρα), their first robotics model derived from the Phi series of vision-language models. Rho-alpha translates natural language commands into control signals for robotic systems performing bimanual manipulation tasks. Commands like "push the green button with the right gripper," "pull out the red wire," "flip the top switch on," or "turn the knob to position 5" get executed directly by dual-arm robots. What makes this different from standard vision-language-action (VLA) models is the additional modalities. Rho-alpha is a VLA+ model that adds tactile sensing to the perceptual mix, with plans to incorporate force feedback. On the learning side, the model is designed to continually improve during deployment by learning from human feedback. The training approach combines trajectories from physical demonstrations and simulated tasks with web-scale visual question answering data. Since teleoperation data is scarce and expensive, Microsoft is using NVIDIA Isaac Sim on Azure to generate physically accurate synthetic datasets via reinforcement learning. These simulated trajectories get combined with commercial and open physical demonstration datasets. The model is currently under evaluation on dual-arm setups and humanoid robots. Microsoft is opening an Early Access Program for organizations interested in evaluating Rho-alpha. Robots that can adapt to dynamic situations and human preferences are more useful in real environments and more trusted by the people operating them. Read more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

61,026 Aufrufe • vor 8 Monaten

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 Aufrufe • vor 9 Monaten