Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. 7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure. High success rate against physical disturbances, zero-shot, and few-shot adaptation Our first step toward a deployable robot learning system.

92,064 görüntüleme • 11 ay önce •via X (Twitter)

30 Yorum

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

If you're troubled by insomnia, you can come and count for this demo.🤣🤣🤣 Task Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Human vs. Robots: Guess who will win?

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

RL-100 introduces a three-stage pipeline: (i) Diffusion policy to acquire human priors; (ii) Iterative offline RL, PPO-style objective applied in the denoising process; to deliver conservative, near-monotonic improvements; (iii) Online RL to eliminate residual failure modes.

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Steer policy to 100% success after 210 trials.

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Takeaway: 1) Variance clipping is valid for stable exploration - variance clipping in the stochastic DDIM sampling process. 2) Epsilon prediction is more suitable for RL: large noise schedule for exploration 3) Reconstruction is crucial for visual robotic manipulation RL as it mitigates representational drift and improves sample efficiency. 4) On a relatively clean scene, the 3D variant learns faster and attains a higher final success rate. 5) CM effectively compresses the iterative denoising process without sacrificing control quality, enabling high-frequency deployment.

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Execution Efficiency: 1) CM (RL) > DDIM (RL) > DP3 (IL) > DP (IL); 2) RL > Human teleoperation

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

High-frequency control is crucial for real robot deployment. Achieving a policy inference frequency of about 378Hz.

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

A huge thank you to my incredible team for all the hard work, dedication, and teamwork. RL-100 team: @kunlei15; @imhuanyuli; @manutd_moon; Zhenyu Wei; @Lingxiao234; @Zhennanjia13007; Ziyu Wang; Shiyu Liang; @HarryXu12 Paper link:

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

@manutd_moon @Lingxiao234 @ZhennanJia13007 @HarryXu12 Guess what the robot has learned: What does it think?

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Achieving a 100% success rate over 7 tasks. If you're troubled by insomnia, you can come and count for this demo.🤣 Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Ted Xiao profil fotoğrafı
Ted Xiao11 ay önce

900/900 is a huge flex! Congrats on the strong results

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Thanks, Ted!

lil’km profil fotoğrafı
lil’km11 ay önce

Congrats Kun! Impressive results 🦾

Théo profil fotoğrafı
Théo11 ay önce

How is the lighting change robustness?

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

In real robot settings, we use a 3d point cloud as visual input without RGB. Thus, there will be no impact on performance from the changed lighting.

Alexander Soare profil fotoğrafı
Alexander Soare11 ay önce

This is amazing work. I love that the experiments section has an explicit focus on long-running uninterrupted success. Geeking out on the specifics here: I'm wondering about the choice of consistency distillation loss, particularly, running the teacher's full denoising chain rather than taking a single step as in the original CM paper.

Alan Zhuolun Zhao profil fotoğrafı
Alan Zhuolun Zhao11 ay önce

Very impressive results! Any reason for the different input/output action spaces? (Joint position vs delta EEF pose)

zkComRi.ETH profil fotoğrafı
zkComRi.ETH11 ay önce

RL-100 turns bots pro

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Thanks!

pfung profil fotoğrafı
pfung11 ay önce

Great results, big congrats! 1. I was curious to see the fail cases in the few shot/one shot trials, are those in the embedded videos? (Didn't want to watch the whole vid to find 😄) 2. Any plans on releasing code for this? Thanks!!!

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

1. No problem, we will release the failure cases. 2. We will release it in the future, after the paper's acceptance.

pfung profil fotoğrafı
pfung11 ay önce

Thanks so much!!

pfung profil fotoğrafı
pfung10 ay önce

1 month ping :) wondering if you still plan to open source RL-100 code. In any case,congrats on a cool paper!

Antoni profil fotoğrafı
Antoni2 ay önce

My take is that the whole trick is how to have a generic critic and how to reset the environment. Once you figure those two out, you can continuously improve on hundreds of tasks

Chi Chu profil fotoğrafı
Chi Chu11 ay önce

100% success rate, that's unbelievable!🤯 Congrats Kun!

Kun Lei profil fotoğrafı
Kun Lei11 ay önce

Thanks, Chi!

Crypto Kɑrhen profil fotoğrafı
Crypto Kɑrhen11 ay önce

insane uptime

Dominique Paul profil fotoğrafı
Dominique Paul10 ay önce

Why did you choose not to release the code for the paper?

sunny profil fotoğrafı
sunny11 ay önce

Great work! When would you publish the code?

MCS @ Safe Sentinel profil fotoğrafı
MCS @ Safe Sentinel11 ay önce

Congrats! Best thing I have seen this month. How does success rate compare on long horizon tasks vs shorter tasks? Hil-serl was bad at long horizon (>9a) but this seems to be way better

Benzer Videolar