Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

So we did a bunch of projects with real world reinforcement learning - but it was often too inefficient to be practical to train tabula rasa. This suggests we need better priors, but acquiring these from on-robot data can often be expensive as well. In our recent work, we...

13,637 Aufrufe • vor 1 Jahr •via X (Twitter)

13 Kommentare

Profilbild von Abhishek Gupta
Abhishek Guptavor 1 Jahr

Ok so SGFT does more than just provide a simulation initialization. Why does this matter? - Standard finetuning methods from simulation can suffer from catastrophic forgetting where performance collapses. SGFT makes consistent rapid progress during finetuning by using simulation to provide persistent guidance through training. SGFT is significantly more sample efficient and can solve difficult contact-rich tasks in the real world where prior methods can fail entirely, often within just minutes of real-world data collection! (2/6)

Profilbild von Abhishek Gupta
Abhishek Guptavor 1 Jahr

So what’s the key idea: while policies may not transfer directly from sim2real due to dynamics mismatch, value functions in simulation capture the approximate geometry of the problem that *does* transfer approximately from sim2real. The ordering of states defined by a sim-learned value function (V_sim) captures successful behaviors that are invariant between sim and real, even if the low-level dynamics differ somewhat. SGFT uses this insight to accelerate real-world finetuning by *using V_sim to perform potential-based reward shaping for real-world RL*. We show both theoretically and empirically that doing so effectively shortens the learning horizon, making learning far more efficient! (3/6)

Profilbild von Abhishek Gupta
Abhishek Guptavor 1 Jahr

SGFT is extremely simple, and easy to integrate with *any* RL-based finetuning paradigm - pick your favorite RL algorithm and add this on top. In particular, SGFT integrates nicely with Model-Based RL (MBRL). Because we shorten the learning horizon, we sidestep the core challenge facing MBRL — prediction errors that compound over long horizons. This enables us to use generative world models such as dreamer or TD-MPC to boost sample efficiency. (4/6)

Profilbild von Abhishek Gupta
Abhishek Guptavor 1 Jahr

We applied this on a number of contact-rich manipulation tasks in the real-world with a Franka robot. Here is a comparison of time to learn each task with our method vs existing baselines using sim2real transfer, RL finetuning, and/or model-based RL. SGFT tends to add value to *any* RL finetuning methods it’s used with, doing a lot better than finetuning methods that simply transfer policy initializations. In each case, our method outperforms baselines in sample efficiency by at least 2x, often learning in just minutes! Some tasks like hammering seem easy, but require getting dynamics precisely right. (5/6)

Profilbild von Abhishek Gupta
Abhishek Guptavor 1 Jahr

The key thing I took away from here is - simulation is inherently wrong, but can still be very useful! Value functions from simulation can make the job of real-world RL *much* easier, making it far more practical as a solution. This was work conceptualized and led by @patrickhyin and Tyler Westenbroek, along with a great set of collaborators - Simran Bagaria, Kevin Huang, @chinganc_rl, @Andrey__Kolobov between UW and MSR Website: Paper: We will be presenting this paper at #ICLR2025 this April 😃

Profilbild von Yang
Yangvor 1 Jahr

Want to learn how practical AI skills and automations for your business and work? Check out our 50+ step-by-step video tutorials 100% FREE 20+ hours of Ai and Automation goodness absolutely free 🥳

Profilbild von Dan Roy 🇨🇦🇩🇰🇱🇹
Dan Roy 🇨🇦🇩🇰🇱🇹vor 1 Jahr

Nice idea. Reminds me of "luckiness" in statistical learning theory. @bremen79

Profilbild von Omar CrazyTechGuy
Omar CrazyTechGuyvor 1 Jahr

Curious to see how you tackled the cost issue in your recent work. Better priors are definitely needed for real-world RL!

Profilbild von Haitham Bou Ammar
Haitham Bou Ammarvor 1 Jahr

Very very interesting; it might be worth checking this to see if it could also further help:

Profilbild von Sergey Levine
Sergey Levinevor 1 Jahr

We can scale up test-time compute with verifiers or without verifiers. It's not obvious which way is better. Turns out we can do some theoretical analysis and show that the use of verifiers is more optimal, and this is supported by experiments!

Profilbild von Minghuan Liu
Minghuan Liuvor 1 Jahr

Thrilled to introduce 🦏RHINO: Learning Real-Time Humanoid-Human-Object Interaction from Human Demonstrations! Project: RHINO is our recent attempt on human-robot interaction to bring humanoids into human's real life. To do so, we require highly responsive interaction abilities for humanoid robots. Instead of multi-stage and language-driven interaction, RHINO supports real-time interruption and automatically responds by recognizing human behaviors! In addition, we integrate a large skill set and make sure RHINO is safe for humans. It's note-worthy that besides object manipulation skills that are learned from teleoperation data, RHINO learns other interaction behaviors from only human-human interaction data, which is much more accessible and scalable! We open-source all training and real-robot deployment codes, along with the datasets for the community! As you may observe, the H1 robot currently only responds using upper-body joints. But I want to highlight that, this is a parallel project to HugWBC ( -- our recent progress on humanoid locomotion, which supports arbitrary upper-body interruption! After we extend the humanoid's ability from both aspects, making humanoids work as real human-like assistants in whole-body is not a distant dream anymore! This work is equally contributed by @ChenTimer, @lixinyao442219, and Jiahang Cao, and is co-worked by @zhu990729, Wentao Dong, me, Prof. Ying Wen, Prof. Yong Yu, Prof. Liqing Zhang, and Prof. Weinan Zhang.

Profilbild von Glen Berseth
Glen Bersethvor 1 Jahr

My next set of lectures on foundational models for #robotics has been posted. These lectures cover learning models, using foundational models to compute rewards, and #scalingdatacollection for robotics.

Profilbild von C Zhang
C Zhangvor 1 Jahr

crazy... G1 has made different labs compete and rush for highly similar topics

Ähnliche Videos

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 Aufrufe • vor 7 Monaten