Loading video...

Video Failed to Load

Go Home

๐Ÿค– Robotics often faces a chicken and egg problem: no web-scale robot data for training (unlike CV or NLP) b/c robots aren't deployed yet & vice-versa. Introducing VRB: Use large-scale human videos to train a *general-purpose* affordance model to jumpstart any robotics paradigm!

70,712 views โ€ข 3 years ago โ€ขvia X (Twitter)

10 Comments

Deepak Pathak's profile picture
Deepak Pathak3 years ago

๐ŸŒ Given a new scene, our general-purpose VRB model predicts all the locations *where* a robot can manipulate objects and *how* should it move post-manipulation. ๐Ÿ’กQs: 1) What are affordances? 2) How to get large-scale data? 3) How does this help robots?

Deepak Pathak's profile picture
Deepak Pathak3 years ago

๐Ÿงฉ In computer vision, there is a long line of work in affordances (Gibson 1966, 1979). However, what's the best way to define them for robotics? Our solution: interaction points & post-contact trajectories, allowing for seamless integration with almost any robot learning setup.

Deepak Pathak's profile picture
Deepak Pathak3 years ago

๐Ÿ”ง๐Ÿ”ง How do we extract affordances from humans? We estimate hand poses & hand-object interaction points to extract contact pts + interaction direction (thanks to @DandanShan_, David Fouhey's 100DoH), and map them back to frames where humans arent present to avoid embodiment gap.

Deepak Pathak's profile picture
Deepak Pathak3 years ago

How to use this model for different robot learning paradigms? We show 4 different paradigms. 1. Offline Data Collection ๐Ÿ“Š Instead of using teleoperation or scripted policies, VRB can allow for an automatic collection of good quality interaction-rich data for policy learning.

Deepak Pathak's profile picture
Deepak Pathak3 years ago

2. Bootstrapping Exploration ๐Ÿ” Robots can utilize the predicted affordances from VRB to explore their environment more intelligently, leading to the discovery of novel and effective ways to manipulate objects.

Deepak Pathak's profile picture
Deepak Pathak3 years ago

3. Goal-Conditioned Learning ๐ŸŽฏ Can we use our affordance model to iteratively improve performance? We train goal-oriented policies using actions sampled from VRB. We prune this action distribution using the distance to the goal during the training process (using WHIRL RSS'22).

Deepak Pathak's profile picture
Deepak Pathak3 years ago

4. Action Space Reparameterization ๐Ÿ”„ We can also treat our affordances as an action space itself. It helps accelerate online RL eliminating the need for task-specific primitives. This enables DQN to learn manipulation tasks in under 30 mins!

Deepak Pathak's profile picture
Deepak Pathak3 years ago

Using internet-scale videos is a promising approach to tackling the data problem in robotics and we hope VRB is a strong step toward that! Led by @shikharbahl, @mendonca_rl, @lchen915, @unnatjain2010 CVPR 2023 Paper: Website: 9/9

Murtaza Dalal's profile picture
Murtaza Dalal3 years ago

Awesome work by my colleagues @shikharbahl, @mendonca_rl, @lchen915! Extracting affordances from human video enables robots to efficiently solve a wide array of manipulation tasks in the real world ๐Ÿค–. Excited to see where this goes next!

Siyuan HUANG's profile picture
Siyuan HUANG3 years ago

Great ideas and a very interesting presentation. :-) I have checked your website but find the dataset is coming soon, would that be available soon?

Related Videos

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's ๐Ÿ”ฅ Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. โ†’ 1,000 hours pre-training โ†’ 20% normalised task performance โ†’ 10,000 hours โ†’ 28% โ†’ 100,000 hours โ†’ 45% โ†’ 1,000,000 hours โ†’ 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ โ™ป๏ธ Join the weekly robotics newsletter, and never miss any news โ†’

Lukas Ziegler

23,576 views โ€ข 1 month ago