Loading video...

Video Failed to Load

Go Home

Our CoRL 2024 paper shows Reinforcement Learning can allow robots to learn skills via real-world practice, without any demonstrations or simulation engineering. Rewards are provided using language/vision models, and mobility of robots enables autonomous exploration. 1/N

38,454 views • 1 year ago •via X (Twitter)

10 Comments

Russell Mendonca's profile picture
Russell Mendonca1 year ago

This timelapse demonstrates the learning process, for moving the chair alternatively between the goals shown. We compare the behavior of the initial policy, to that of a later policy after about 8-10 hours of autonomous practice. 2/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

Exploration for complex robots is hard due to the very large action space, and since actions rarely provide useful signal for the task, as seen on the left. Instead, we encourage the robots to grasp objects before exploring how to perform tasks, seen on the right. 3/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

After making some progress on the task, the robot should not stagnate near goal states. We use goal cycles for mobile systems, where one goal serves to reset the other. This allows the robot to preserve state diversity. 4/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

To use this signal rich data to learn skills, we use RL. For greater sample efficiency, we combine RL with behavior priors that contain basic task knowledge. These priors can be planners with a simplified incomplete model, or procedurally generated motions. 5/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

For rewards, we use detection and segmentation models, which given the name of an object of interest, produce the corresponding mask. Combined with low-level depth observations, we get a state estimate of the object. This is compared to the desired goal state for reward. 6/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

Hence our framework consists of 3 main pieces - 1) Task-relevant autonomy, which ensures data collected is likely to have learning signal, 2) Efficient Control, which seeks to learn skills quickly using this signal, 3) Flexible supervision, for defining the learning signal. 7/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

We find task-relevant autonomy is critical for any meaningful progress. Here we compare our approach of efficient control to using only the behavior prior, or RL without the prior (all use task-relevant autonomy). The former is insufficient and the latter learns very slowly. 8/N

Russell Mendonca's profile picture
Russell Mendonca1 year ago

Thanks to the AI Institute for this collaboration, and my co-authors Emmanuel Panov, @b_k_bucher, Jiuguang Wang and @pathak2206 ! Please see the website ( and paper ( for more details. 9/9

Kyle Stachowicz's profile picture
Kyle Stachowicz1 year ago

Nice work, these look like some pretty challenging tasks! Cool to see more people doing online RL directly in the real world :)

Kevin Zakka's profile picture
Kevin Zakka1 year ago

Congrats!

Related Videos

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,533 views • 6 months ago