Загрузка видео...

Не удалось загрузить видео

На главную

Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. 7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure. High success rate against physical disturbances, zero-shot, and few-shot adaptation Our first step toward a deployable robot learning system.

92,064 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 30

Фото профиля Kun Lei
Kun Lei11 месяцев назад

If you're troubled by insomnia, you can come and count for this demo.🤣🤣🤣 Task Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Human vs. Robots: Guess who will win?

Фото профиля Kun Lei
Kun Lei11 месяцев назад

RL-100 introduces a three-stage pipeline: (i) Diffusion policy to acquire human priors; (ii) Iterative offline RL, PPO-style objective applied in the denoising process; to deliver conservative, near-monotonic improvements; (iii) Online RL to eliminate residual failure modes.

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Steer policy to 100% success after 210 trials.

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Takeaway: 1) Variance clipping is valid for stable exploration - variance clipping in the stochastic DDIM sampling process. 2) Epsilon prediction is more suitable for RL: large noise schedule for exploration 3) Reconstruction is crucial for visual robotic manipulation RL as it mitigates representational drift and improves sample efficiency. 4) On a relatively clean scene, the 3D variant learns faster and attains a higher final success rate. 5) CM effectively compresses the iterative denoising process without sacrificing control quality, enabling high-frequency deployment.

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Execution Efficiency: 1) CM (RL) > DDIM (RL) > DP3 (IL) > DP (IL); 2) RL > Human teleoperation

Фото профиля Kun Lei
Kun Lei11 месяцев назад

High-frequency control is crucial for real robot deployment. Achieving a policy inference frequency of about 378Hz.

Фото профиля Kun Lei
Kun Lei11 месяцев назад

A huge thank you to my incredible team for all the hard work, dedication, and teamwork. RL-100 team: @kunlei15; @imhuanyuli; @manutd_moon; Zhenyu Wei; @Lingxiao234; @Zhennanjia13007; Ziyu Wang; Shiyu Liang; @HarryXu12 Paper link:

Фото профиля Kun Lei
Kun Lei11 месяцев назад

@manutd_moon @Lingxiao234 @ZhennanJia13007 @HarryXu12 Guess what the robot has learned: What does it think?

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Achieving a 100% success rate over 7 tasks. If you're troubled by insomnia, you can come and count for this demo.🤣 Timeline: Folding: 05:18-16:09 Juicing-placing: 16:09-19:09 Juicing-removal: 19:09-21:21 Pouring: 21:21-23:12 Unscrewing: 23:12-24:57 Bowling: 24:57-27:26 Push-T: 27:26-30:21 OverallJuicing: 30:21-31:57

Фото профиля Ted Xiao
Ted Xiao11 месяцев назад

900/900 is a huge flex! Congrats on the strong results

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Thanks, Ted!

Фото профиля lil’km
lil’km11 месяцев назад

Congrats Kun! Impressive results 🦾

Фото профиля Théo
Théo11 месяцев назад

How is the lighting change robustness?

Фото профиля Kun Lei
Kun Lei11 месяцев назад

In real robot settings, we use a 3d point cloud as visual input without RGB. Thus, there will be no impact on performance from the changed lighting.

Фото профиля Alexander Soare
Alexander Soare11 месяцев назад

This is amazing work. I love that the experiments section has an explicit focus on long-running uninterrupted success. Geeking out on the specifics here: I'm wondering about the choice of consistency distillation loss, particularly, running the teacher's full denoising chain rather than taking a single step as in the original CM paper.

Фото профиля Alan Zhuolun Zhao
Alan Zhuolun Zhao11 месяцев назад

Very impressive results! Any reason for the different input/output action spaces? (Joint position vs delta EEF pose)

Фото профиля zkComRi.ETH
zkComRi.ETH11 месяцев назад

RL-100 turns bots pro

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Thanks!

Фото профиля pfung
pfung11 месяцев назад

Great results, big congrats! 1. I was curious to see the fail cases in the few shot/one shot trials, are those in the embedded videos? (Didn't want to watch the whole vid to find 😄) 2. Any plans on releasing code for this? Thanks!!!

Фото профиля Kun Lei
Kun Lei11 месяцев назад

1. No problem, we will release the failure cases. 2. We will release it in the future, after the paper's acceptance.

Фото профиля pfung
pfung11 месяцев назад

Thanks so much!!

Фото профиля pfung
pfung10 месяцев назад

1 month ping :) wondering if you still plan to open source RL-100 code. In any case,congrats on a cool paper!

Фото профиля Antoni
Antoni2 месяцев назад

My take is that the whole trick is how to have a generic critic and how to reset the environment. Once you figure those two out, you can continuously improve on hundreds of tasks

Фото профиля Chi Chu
Chi Chu11 месяцев назад

100% success rate, that's unbelievable!🤯 Congrats Kun!

Фото профиля Kun Lei
Kun Lei11 месяцев назад

Thanks, Chi!

Фото профиля Crypto Kɑrhen
Crypto Kɑrhen11 месяцев назад

insane uptime

Фото профиля Dominique Paul
Dominique Paul10 месяцев назад

Why did you choose not to release the code for the paper?

Фото профиля sunny
sunny11 месяцев назад

Great work! When would you publish the code?

Фото профиля MCS @ Safe Sentinel
MCS @ Safe Sentinel11 месяцев назад

Congrats! Best thing I have seen this month. How does success rate compare on long horizon tasks vs shorter tasks? Hil-serl was bad at long horizon (>9a) but this seems to be way better

Похожие видео