Загрузка видео...

Не удалось загрузить видео

На главную

Wouldn't it be great if we could train robots without any teleoperation! In our latest paper, we train robots to mimic a human video of the task by simply matching the object features using RL. We only need one video and under an hour of robot training.

46,221 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля Lerrel Pinto
Lerrel Pinto1 год назад

HuDOR generates rewards from human videos by tracking points on manipulable objects using off-the-shelf trackers. Then during online learning, the robot optimizes its policy to match the trajectory of the manipulated objects to those in the human video.

Фото профиля Lerrel Pinto
Lerrel Pinto1 год назад

HuDOR also generalizes to new objects! Instead of training policies from scratch, we use different text prompts to get object point movements and apply pre-trained HuDOR policies. HuDOR reasonably generalizes, with failures mainly due to weight.

Фото профиля Lerrel Pinto
Lerrel Pinto1 год назад

To generalize to larger areas, we decouple reaching and manipulation. Using object detection and relative transforms, we locate objects in both robot and human demos, calculate offsets, and integrate them into HuDOR-trained policies, achieving ~60% success within a 40x50 cm area.

Фото профиля Lerrel Pinto
Lerrel Pinto1 год назад

This work was led by @irmakkguzey with @wxdyl0915, @georgysavva and @Raunaqmb. For more details, including additional robot videos, visit our website:

Фото профиля Homanga Bharadhwaj
Homanga Bharadhwaj1 год назад

This is very cool work, Lerrel. And congrats @irmakkguzey ! I was thinking about something very similar recently, doing real-world RL based on a human demo in the same scene (and was concerned about resets) It's great to see you have validated this already :)

Фото профиля Thanos Variant
Thanos Variant1 год назад

I wonder if robots will ever be able to improvise if you remove their hands. Can they still perform the task in a different way? How would you even train that?🤔

Фото профиля Sergey Golubev
Sergey Golubev1 год назад

Impressive advancement! Using object feature matching and RL to train robots without teleoperation opens new possibilities in human-robot interaction.

Фото профиля Crynet
Crynet1 год назад

Training robots from a single human video using RL could revolutionize robotics. Reducing data needs and training time is a significant step towards more adaptable AI systems.

Фото профиля Dave Robertson
Dave Robertson1 год назад

@ylecun Too slow

Фото профиля Chris Thompson
Chris Thompson1 год назад

Love the idea of ditching teleoperation! What's the most challenging part of implementing this in real-world scenarios?

Похожие видео

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

293,585 просмотров • 5 месяцев назад