Loading video...

Video Failed to Load

Go Home

Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. 7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure. High success rate against physical disturbances, zero-shot, and few-shot adaptation Our first step toward a deployable robot learning system.

92,064 views • 11 months ago •via X (Twitter)

30 Comments

Kun Lei's profile picture
Kun Lei11 months ago

If you're troubled by insomnia, you can come and count for this demo.🤣🤣🤣 Task Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Kun Lei's profile picture
Kun Lei11 months ago

Human vs. Robots: Guess who will win?

Kun Lei's profile picture
Kun Lei11 months ago

RL-100 introduces a three-stage pipeline: (i) Diffusion policy to acquire human priors; (ii) Iterative offline RL, PPO-style objective applied in the denoising process; to deliver conservative, near-monotonic improvements; (iii) Online RL to eliminate residual failure modes.

Kun Lei's profile picture
Kun Lei11 months ago

Steer policy to 100% success after 210 trials.

Kun Lei's profile picture
Kun Lei11 months ago

Takeaway: 1) Variance clipping is valid for stable exploration - variance clipping in the stochastic DDIM sampling process. 2) Epsilon prediction is more suitable for RL: large noise schedule for exploration 3) Reconstruction is crucial for visual robotic manipulation RL as it mitigates representational drift and improves sample efficiency. 4) On a relatively clean scene, the 3D variant learns faster and attains a higher final success rate. 5) CM effectively compresses the iterative denoising process without sacrificing control quality, enabling high-frequency deployment.

Kun Lei's profile picture
Kun Lei11 months ago

Execution Efficiency: 1) CM (RL) > DDIM (RL) > DP3 (IL) > DP (IL); 2) RL > Human teleoperation

Kun Lei's profile picture
Kun Lei11 months ago

High-frequency control is crucial for real robot deployment. Achieving a policy inference frequency of about 378Hz.

Kun Lei's profile picture
Kun Lei11 months ago

A huge thank you to my incredible team for all the hard work, dedication, and teamwork. RL-100 team: @kunlei15; @imhuanyuli; @manutd_moon; Zhenyu Wei; @Lingxiao234; @Zhennanjia13007; Ziyu Wang; Shiyu Liang; @HarryXu12 Paper link:

Kun Lei's profile picture
Kun Lei11 months ago

@manutd_moon @Lingxiao234 @ZhennanJia13007 @HarryXu12 Guess what the robot has learned: What does it think?

Kun Lei's profile picture
Kun Lei11 months ago

Achieving a 100% success rate over 7 tasks. If you're troubled by insomnia, you can come and count for this demo.🤣 Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Ted Xiao's profile picture
Ted Xiao11 months ago

900/900 is a huge flex! Congrats on the strong results

Kun Lei's profile picture
Kun Lei11 months ago

Thanks, Ted!

lil’km's profile picture
lil’km11 months ago

Congrats Kun! Impressive results 🦾

Théo's profile picture
Théo11 months ago

How is the lighting change robustness?

Kun Lei's profile picture
Kun Lei11 months ago

In real robot settings, we use a 3d point cloud as visual input without RGB. Thus, there will be no impact on performance from the changed lighting.

Alexander Soare's profile picture
Alexander Soare11 months ago

This is amazing work. I love that the experiments section has an explicit focus on long-running uninterrupted success. Geeking out on the specifics here: I'm wondering about the choice of consistency distillation loss, particularly, running the teacher's full denoising chain rather than taking a single step as in the original CM paper.

Alan Zhuolun Zhao's profile picture
Alan Zhuolun Zhao11 months ago

Very impressive results! Any reason for the different input/output action spaces? (Joint position vs delta EEF pose)

zkComRi.ETH's profile picture
zkComRi.ETH11 months ago

RL-100 turns bots pro

Kun Lei's profile picture
Kun Lei11 months ago

Thanks!

pfung's profile picture
pfung11 months ago

Great results, big congrats! 1. I was curious to see the fail cases in the few shot/one shot trials, are those in the embedded videos? (Didn't want to watch the whole vid to find 😄) 2. Any plans on releasing code for this? Thanks!!!

Kun Lei's profile picture
Kun Lei11 months ago

1. No problem, we will release the failure cases. 2. We will release it in the future, after the paper's acceptance.

pfung's profile picture
pfung11 months ago

Thanks so much!!

pfung's profile picture
pfung10 months ago

1 month ping :) wondering if you still plan to open source RL-100 code. In any case,congrats on a cool paper!

Antoni's profile picture
Antoni2 months ago

My take is that the whole trick is how to have a generic critic and how to reset the environment. Once you figure those two out, you can continuously improve on hundreds of tasks

Chi Chu's profile picture
Chi Chu11 months ago

100% success rate, that's unbelievable!🤯 Congrats Kun!

Kun Lei's profile picture
Kun Lei11 months ago

Thanks, Chi!

Crypto Kɑrhen's profile picture
Crypto Kɑrhen11 months ago

insane uptime

Dominique Paul's profile picture
Dominique Paul10 months ago

Why did you choose not to release the code for the paper?

sunny's profile picture
sunny11 months ago

Great work! When would you publish the code?

MCS @ Safe Sentinel's profile picture
MCS @ Safe Sentinel11 months ago

Congrats! Best thing I have seen this month. How does success rate compare on long horizon tasks vs shorter tasks? Hil-serl was bad at long horizon (>9a) but this seems to be way better

Related Videos