Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🤖 Robotics often faces a chicken and egg problem: no web-scale robot data for training (unlike CV or NLP) b/c robots aren't deployed yet & vice-versa. Introducing VRB: Use large-scale human videos to train a *general-purpose* affordance model to jumpstart any robotics paradigm!

70,712 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

🌐 Given a new scene, our general-purpose VRB model predicts all the locations *where* a robot can manipulate objects and *how* should it move post-manipulation. 💡Qs: 1) What are affordances? 2) How to get large-scale data? 3) How does this help robots?

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

🧩 In computer vision, there is a long line of work in affordances (Gibson 1966, 1979). However, what's the best way to define them for robotics? Our solution: interaction points & post-contact trajectories, allowing for seamless integration with almost any robot learning setup.

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

🔧🔧 How do we extract affordances from humans? We estimate hand poses & hand-object interaction points to extract contact pts + interaction direction (thanks to @DandanShan_, David Fouhey's 100DoH), and map them back to frames where humans arent present to avoid embodiment gap.

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

How to use this model for different robot learning paradigms? We show 4 different paradigms. 1. Offline Data Collection 📊 Instead of using teleoperation or scripted policies, VRB can allow for an automatic collection of good quality interaction-rich data for policy learning.

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

2. Bootstrapping Exploration 🔍 Robots can utilize the predicted affordances from VRB to explore their environment more intelligently, leading to the discovery of novel and effective ways to manipulate objects.

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

3. Goal-Conditioned Learning 🎯 Can we use our affordance model to iteratively improve performance? We train goal-oriented policies using actions sampled from VRB. We prune this action distribution using the distance to the goal during the training process (using WHIRL RSS'22).

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

4. Action Space Reparameterization 🔄 We can also treat our affordances as an action space itself. It helps accelerate online RL eliminating the need for task-specific primitives. This enables DQN to learn manipulation tasks in under 30 mins!

Profilbild von Deepak Pathak
Deepak Pathakvor 3 Jahren

Using internet-scale videos is a promising approach to tackling the data problem in robotics and we hope VRB is a strong step toward that! Led by @shikharbahl, @mendonca_rl, @lchen915, @unnatjain2010 CVPR 2023 Paper: Website: 9/9

Profilbild von Murtaza Dalal
Murtaza Dalalvor 3 Jahren

Awesome work by my colleagues @shikharbahl, @mendonca_rl, @lchen915! Extracting affordances from human video enables robots to efficiently solve a wide array of manipulation tasks in the real world 🤖. Excited to see where this goes next!

Profilbild von Siyuan HUANG
Siyuan HUANGvor 3 Jahren

Great ideas and a very interesting presentation. :-) I have checked your website but find the dataset is coming soon, would that be available soon?

Ähnliche Videos

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,576 Aufrufe • vor 1 Monat