Video yükleniyor...
Video Yüklenemedi
Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. 7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure. High success rate against physical disturbances, zero-shot, and few-shot adaptation Our first step toward a deployable robot learning system.
92,064 görüntüleme • 11 ay önce •via X (Twitter)
30 Yorum

If you're troubled by insomnia, you can come and count for this demo.🤣🤣🤣 Task Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Human vs. Robots: Guess who will win?

RL-100 introduces a three-stage pipeline: (i) Diffusion policy to acquire human priors; (ii) Iterative offline RL, PPO-style objective applied in the denoising process; to deliver conservative, near-monotonic improvements; (iii) Online RL to eliminate residual failure modes.

Steer policy to 100% success after 210 trials.

Takeaway: 1) Variance clipping is valid for stable exploration - variance clipping in the stochastic DDIM sampling process. 2) Epsilon prediction is more suitable for RL: large noise schedule for exploration 3) Reconstruction is crucial for visual robotic manipulation RL as it mitigates representational drift and improves sample efficiency. 4) On a relatively clean scene, the 3D variant learns faster and attains a higher final success rate. 5) CM effectively compresses the iterative denoising process without sacrificing control quality, enabling high-frequency deployment.

Execution Efficiency: 1) CM (RL) > DDIM (RL) > DP3 (IL) > DP (IL); 2) RL > Human teleoperation

High-frequency control is crucial for real robot deployment. Achieving a policy inference frequency of about 378Hz.

A huge thank you to my incredible team for all the hard work, dedication, and teamwork. RL-100 team: @kunlei15; @imhuanyuli; @manutd_moon; Zhenyu Wei; @Lingxiao234; @Zhennanjia13007; Ziyu Wang; Shiyu Liang; @HarryXu12 Paper link:

@manutd_moon @Lingxiao234 @ZhennanJia13007 @HarryXu12 Guess what the robot has learned: What does it think?

Achieving a 100% success rate over 7 tasks. If you're troubled by insomnia, you can come and count for this demo.🤣 Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

900/900 is a huge flex! Congrats on the strong results

Thanks, Ted!

Congrats Kun! Impressive results 🦾

How is the lighting change robustness?

In real robot settings, we use a 3d point cloud as visual input without RGB. Thus, there will be no impact on performance from the changed lighting.

This is amazing work. I love that the experiments section has an explicit focus on long-running uninterrupted success. Geeking out on the specifics here: I'm wondering about the choice of consistency distillation loss, particularly, running the teacher's full denoising chain rather than taking a single step as in the original CM paper.

Very impressive results! Any reason for the different input/output action spaces? (Joint position vs delta EEF pose)

RL-100 turns bots pro

Thanks!

Great results, big congrats! 1. I was curious to see the fail cases in the few shot/one shot trials, are those in the embedded videos? (Didn't want to watch the whole vid to find 😄) 2. Any plans on releasing code for this? Thanks!!!

1. No problem, we will release the failure cases. 2. We will release it in the future, after the paper's acceptance.

Thanks so much!!

1 month ping :) wondering if you still plan to open source RL-100 code. In any case,congrats on a cool paper!

My take is that the whole trick is how to have a generic critic and how to reset the environment. Once you figure those two out, you can continuously improve on hundreds of tasks

100% success rate, that's unbelievable!🤯 Congrats Kun!

Thanks, Chi!

insane uptime

Why did you choose not to release the code for the paper?

Great work! When would you publish the code?

Congrats! Best thing I have seen this month. How does success rate compare on long horizon tasks vs shorter tasks? Hil-serl was bad at long horizon (>9a) but this seems to be way better
