Video wird geladen...
Video konnte nicht geladen werden
🤖 Robotics often faces a chicken and egg problem: no web-scale robot data for training (unlike CV or NLP) b/c robots aren't deployed yet & vice-versa. Introducing VRB: Use large-scale human videos to train a *general-purpose* affordance model to jumpstart any robotics paradigm!
70,712 Aufrufe • vor 3 Jahren •via X (Twitter)
10 Kommentare

🌐 Given a new scene, our general-purpose VRB model predicts all the locations *where* a robot can manipulate objects and *how* should it move post-manipulation. 💡Qs: 1) What are affordances? 2) How to get large-scale data? 3) How does this help robots?

🧩 In computer vision, there is a long line of work in affordances (Gibson 1966, 1979). However, what's the best way to define them for robotics? Our solution: interaction points & post-contact trajectories, allowing for seamless integration with almost any robot learning setup.

🔧🔧 How do we extract affordances from humans? We estimate hand poses & hand-object interaction points to extract contact pts + interaction direction (thanks to @DandanShan_, David Fouhey's 100DoH), and map them back to frames where humans arent present to avoid embodiment gap.

How to use this model for different robot learning paradigms? We show 4 different paradigms. 1. Offline Data Collection 📊 Instead of using teleoperation or scripted policies, VRB can allow for an automatic collection of good quality interaction-rich data for policy learning.

2. Bootstrapping Exploration 🔍 Robots can utilize the predicted affordances from VRB to explore their environment more intelligently, leading to the discovery of novel and effective ways to manipulate objects.

3. Goal-Conditioned Learning 🎯 Can we use our affordance model to iteratively improve performance? We train goal-oriented policies using actions sampled from VRB. We prune this action distribution using the distance to the goal during the training process (using WHIRL RSS'22).

4. Action Space Reparameterization 🔄 We can also treat our affordances as an action space itself. It helps accelerate online RL eliminating the need for task-specific primitives. This enables DQN to learn manipulation tasks in under 30 mins!

Using internet-scale videos is a promising approach to tackling the data problem in robotics and we hope VRB is a strong step toward that! Led by @shikharbahl, @mendonca_rl, @lchen915, @unnatjain2010 CVPR 2023 Paper: Website: 9/9

Awesome work by my colleagues @shikharbahl, @mendonca_rl, @lchen915! Extracting affordances from human video enables robots to efficiently solve a wide array of manipulation tasks in the real world 🤖. Excited to see where this goes next!

Great ideas and a very interesting presentation. :-) I have checked your website but find the dataset is coming soon, would that be available soon?
