正在加载视频...

视频加载失败

Wouldn't it be great if we could train robots without any teleoperation! In our latest paper, we train robots to mimic a human video of the task by simply matching the object features using RL. We only need one video and under an hour of robot training.

46,221 次观看 • 1 年前 •via X (Twitter)

10 条评论

Lerrel Pinto 的头像
Lerrel Pinto1 年前

HuDOR generates rewards from human videos by tracking points on manipulable objects using off-the-shelf trackers. Then during online learning, the robot optimizes its policy to match the trajectory of the manipulated objects to those in the human video.

Lerrel Pinto 的头像
Lerrel Pinto1 年前

HuDOR also generalizes to new objects! Instead of training policies from scratch, we use different text prompts to get object point movements and apply pre-trained HuDOR policies. HuDOR reasonably generalizes, with failures mainly due to weight.

Lerrel Pinto 的头像
Lerrel Pinto1 年前

To generalize to larger areas, we decouple reaching and manipulation. Using object detection and relative transforms, we locate objects in both robot and human demos, calculate offsets, and integrate them into HuDOR-trained policies, achieving ~60% success within a 40x50 cm area.

Lerrel Pinto 的头像
Lerrel Pinto1 年前

This work was led by @irmakkguzey with @wxdyl0915, @georgysavva and @Raunaqmb. For more details, including additional robot videos, visit our website:

Homanga Bharadhwaj 的头像
Homanga Bharadhwaj1 年前

This is very cool work, Lerrel. And congrats @irmakkguzey ! I was thinking about something very similar recently, doing real-world RL based on a human demo in the same scene (and was concerned about resets) It's great to see you have validated this already :)

Thanos Variant 的头像
Thanos Variant1 年前

I wonder if robots will ever be able to improvise if you remove their hands. Can they still perform the task in a different way? How would you even train that?🤔

Sergey Golubev 的头像
Sergey Golubev1 年前

Impressive advancement! Using object feature matching and RL to train robots without teleoperation opens new possibilities in human-robot interaction.

Crynet 的头像
Crynet1 年前

Training robots from a single human video using RL could revolutionize robotics. Reducing data needs and training time is a significant step towards more adaptable AI systems.

Dave Robertson 的头像
Dave Robertson1 年前

@ylecun Too slow

Chris Thompson 的头像
Chris Thompson1 年前

Love the idea of ditching teleoperation! What's the most challenging part of implementing this in real-world scenarios?

相关视频

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

293,585 次观看 • 5 个月前