Video wird geladen...
Video konnte nicht geladen werden
What if robots could improve themselves by learning from their own failures in the real-world? Introducing 𝗣𝗟𝗗 (𝗣𝗿𝗼𝗯𝗲, 𝗟𝗲𝗮𝗿𝗻, 𝗗𝗶𝘀𝘁𝗶𝗹𝗹) — a recipe that enables Vision-Language-Action (VLA) models to self-improve for high-precision manipulation tasks. PLD couples real-world residual reinforcement learning with standard supervised fine-tuning — letting robots discover, recover,... show more
186,079 Aufrufe • vor 11 Monaten •via X (Twitter)
35 Kommentare

Supervised fine-tuning (SFT) made today’s strongest robot foundation models in applications — but every gain costs fleets of human teleoperators. PLD replaces them with residual RL specialists that only “take over” when the base policy (any VLA architecture) fails. 🧠 Probe the failure → Learn recovery behavior → Distill back into the generalist. And it works: ✅ Minimal human guidance ✅ ~99 % success on LIBERO ✅ +50 % gain in SimplerEnv ✅ 100 % success on real Franka & YAM arms ✅ 1 hr continuous GPU insertion w/o human reset

To perform real-world RL, the algorithm must be super sample-efficient. PLD builds on a series of design choices that make off-policy RL actually practical for robots: 1️⃣ Roll out the base policy longer to flood the demo buffer and oversample successful trials — (RLPD, 2️⃣ Warm up the critic with calibrated Q-learning loss — (Cal-QL, 3️⃣ Carefully choose small initial action scales to prevent RL from sabotaging base policy performance during RL's early stage exploration; 4️⃣ Randomize initial states distribution via base-policy rollouts — unlike JSRL, we skip the scheduler to retain recovery capability (JSRL,

Why does it generalize? Because PLD collects policy-aligned data — not "random" teleop demos. It learns from the exact states the model visits and fails in. A model that trains on its own distribution forgets less and adapts faster. Try the interactive demos on our website to get more understanding👇

Very fortunate to work with this incredible team: @DarthUtopian, @andy_peng05, @HaoruXue, @TairanHe99, @yuqi_xie5, @fengyuan_hu, @jimmyyhwu, @zhengyiluo, @DrJimFan, @GuanyaShi, @yukez 🔗 Check out more demos on our project website: ❤️ We are also deeply grateful for the help and support from @JasonJZLiu, @_tonytao_, @qiyang_li, @letian_fu, @AjayMandlekar, @youliangtan, @Haoyu_Xiong_, @CharlesXu0124, @guanzhi_wang

Is human in the loop correction also possible along with RL corrections?

PLD aims to minimize human guidance. But having human in the loop is definitely possible (e.g., ConRFT,

Impressive results, nice work!

Love this. Robots getting better by failing in the real world feels like the most human way to learn.

The RAM needs to go in Slot 1 and 3, not 1 and 2. OK, then proceed to put it in Slot 3 and 4. 😅

Hi, any plan to open source the project code?

Congrats!!!

This could democratize robotics development. Small teams could deploy basic robots that improve themselves rather than needing massive training infrastructure.

Congrats! What are the key ingredients you observed during real-world RL?

Thanks Minghuan! From my experience, the most critical ingredient for real-world RL is the prior — i.e., the base policy. My take is: If human budget is limited, spend nearly all of it on collect teleop data to pretrain a strong base model. You can see this on my website — there’s a panel visualizing the residual policy output. When the base is good, the residual outputs near-zero most of the time and only corrects at key steps. Early on, I tried a trivial ResNet+MLP base trained with a simple BC loss — RL completely failed, even on basic goal reaching. After switching to ACT and π₀, RL learning became dramatically faster and more stable.

Awesome! Further question: what kind of ability do you expect the base model to cover (for example, which kind of generality)? And what kind of ability does RL deeply improve? Also, is real-world reward design as tough as the ones in training legged robots?

Really cool

That's a fascinating concept, Wenli! Imagine robots actually learning from their mistakes – it's like a whole new level of "jugad", isn't it?

Thanks Himanshu! Exactly, it is all about enabling robots to learn and adapt from their own failures, turning “jugaad” into a systematic self-improvement loop.

Hey I would love to use this for my robotics company let’s connect!

Amazing work!

@apagajewski of interest

This is a very impressive work, but you cant say that you achieved 99% on Libero when you didnt include the numbers for libero long and 90..

Hello, do you plan to release the code?

PLD's 'learn from failure' approach is exactly what warehouse automation needs. Watching robots refine skills in real-time? That's how you build precision. #Automation #Robotics

Good stuff!

@nilscmr

Congrats! Amazing work for real world RL + VLA

very nice

Where can I find the paper?

The arXiv version will probably come out next week

Great work, congrats!

Great work!Would you release the code recently?

@BrianRoemmele

This is fascinating. I’m exploring metaphor-driven wellness tech and emotional UX—curious how PLD handles edge cases where failure is emotionally or physically ambiguous. Can VLA models learn from subtle, non-binary outcomes like hesitation or partial success?

is the project open-sourced? would love to replicate it
