Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. 7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure. High success rate against physical disturbances, zero-shot, and few-shot adaptation Our first step toward a deployable robot learning system.

92,064 Aufrufe • vor 11 Monaten •via X (Twitter)

30 Kommentare

Profilbild von Kun Lei
Kun Leivor 11 Monaten

If you're troubled by insomnia, you can come and count for this demo.🤣🤣🤣 Task Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Human vs. Robots: Guess who will win?

Profilbild von Kun Lei
Kun Leivor 11 Monaten

RL-100 introduces a three-stage pipeline: (i) Diffusion policy to acquire human priors; (ii) Iterative offline RL, PPO-style objective applied in the denoising process; to deliver conservative, near-monotonic improvements; (iii) Online RL to eliminate residual failure modes.

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Steer policy to 100% success after 210 trials.

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Takeaway: 1) Variance clipping is valid for stable exploration - variance clipping in the stochastic DDIM sampling process. 2) Epsilon prediction is more suitable for RL: large noise schedule for exploration 3) Reconstruction is crucial for visual robotic manipulation RL as it mitigates representational drift and improves sample efficiency. 4) On a relatively clean scene, the 3D variant learns faster and attains a higher final success rate. 5) CM effectively compresses the iterative denoising process without sacrificing control quality, enabling high-frequency deployment.

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Execution Efficiency: 1) CM (RL) > DDIM (RL) > DP3 (IL) > DP (IL); 2) RL > Human teleoperation

Profilbild von Kun Lei
Kun Leivor 11 Monaten

High-frequency control is crucial for real robot deployment. Achieving a policy inference frequency of about 378Hz.

Profilbild von Kun Lei
Kun Leivor 11 Monaten

A huge thank you to my incredible team for all the hard work, dedication, and teamwork. RL-100 team: @kunlei15; @imhuanyuli; @manutd_moon; Zhenyu Wei; @Lingxiao234; @Zhennanjia13007; Ziyu Wang; Shiyu Liang; @HarryXu12 Paper link:

Profilbild von Kun Lei
Kun Leivor 11 Monaten

@manutd_moon @Lingxiao234 @ZhennanJia13007 @HarryXu12 Guess what the robot has learned: What does it think?

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Achieving a 100% success rate over 7 tasks. If you're troubled by insomnia, you can come and count for this demo.🤣 Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Profilbild von Ted Xiao
Ted Xiaovor 11 Monaten

900/900 is a huge flex! Congrats on the strong results

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Thanks, Ted!

Profilbild von lil’km
lil’kmvor 11 Monaten

Congrats Kun! Impressive results 🦾

Profilbild von Théo
Théovor 11 Monaten

How is the lighting change robustness?

Profilbild von Kun Lei
Kun Leivor 11 Monaten

In real robot settings, we use a 3d point cloud as visual input without RGB. Thus, there will be no impact on performance from the changed lighting.

Profilbild von Alexander Soare
Alexander Soarevor 11 Monaten

This is amazing work. I love that the experiments section has an explicit focus on long-running uninterrupted success. Geeking out on the specifics here: I'm wondering about the choice of consistency distillation loss, particularly, running the teacher's full denoising chain rather than taking a single step as in the original CM paper.

Profilbild von Alan Zhuolun Zhao
Alan Zhuolun Zhaovor 11 Monaten

Very impressive results! Any reason for the different input/output action spaces? (Joint position vs delta EEF pose)

Profilbild von zkComRi.ETH
zkComRi.ETHvor 11 Monaten

RL-100 turns bots pro

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Thanks!

Profilbild von pfung
pfungvor 11 Monaten

Great results, big congrats! 1. I was curious to see the fail cases in the few shot/one shot trials, are those in the embedded videos? (Didn't want to watch the whole vid to find 😄) 2. Any plans on releasing code for this? Thanks!!!

Profilbild von Kun Lei
Kun Leivor 11 Monaten

1. No problem, we will release the failure cases. 2. We will release it in the future, after the paper's acceptance.

Profilbild von pfung
pfungvor 11 Monaten

Thanks so much!!

Profilbild von pfung
pfungvor 10 Monaten

1 month ping :) wondering if you still plan to open source RL-100 code. In any case,congrats on a cool paper!

Profilbild von Antoni
Antonivor 2 Monaten

My take is that the whole trick is how to have a generic critic and how to reset the environment. Once you figure those two out, you can continuously improve on hundreds of tasks

Profilbild von Chi Chu
Chi Chuvor 11 Monaten

100% success rate, that's unbelievable!🤯 Congrats Kun!

Profilbild von Kun Lei
Kun Leivor 11 Monaten

Thanks, Chi!

Profilbild von Crypto Kɑrhen
Crypto Kɑrhenvor 11 Monaten

insane uptime

Profilbild von Dominique Paul
Dominique Paulvor 10 Monaten

Why did you choose not to release the code for the paper?

Profilbild von sunny
sunnyvor 11 Monaten

Great work! When would you publish the code?

Profilbild von MCS @ Safe Sentinel
MCS @ Safe Sentinelvor 11 Monaten

Congrats! Best thing I have seen this month. How does success rate compare on long horizon tasks vs shorter tasks? Hil-serl was bad at long horizon (>9a) but this seems to be way better

Ähnliche Videos