Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🤖 Robotics often faces a chicken and egg problem: no web-scale robot data for training (unlike CV or NLP) b/c robots aren't deployed yet & vice-versa. Introducing VRB: Use large-scale human videos to train a *general-purpose* affordance model to jumpstart any robotics paradigm!

70,712 görüntüleme • 3 yıl önce •via X (Twitter)

10 Yorum

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

🌐 Given a new scene, our general-purpose VRB model predicts all the locations *where* a robot can manipulate objects and *how* should it move post-manipulation. 💡Qs: 1) What are affordances? 2) How to get large-scale data? 3) How does this help robots?

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

🧩 In computer vision, there is a long line of work in affordances (Gibson 1966, 1979). However, what's the best way to define them for robotics? Our solution: interaction points & post-contact trajectories, allowing for seamless integration with almost any robot learning setup.

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

🔧🔧 How do we extract affordances from humans? We estimate hand poses & hand-object interaction points to extract contact pts + interaction direction (thanks to @DandanShan_, David Fouhey's 100DoH), and map them back to frames where humans arent present to avoid embodiment gap.

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

How to use this model for different robot learning paradigms? We show 4 different paradigms. 1. Offline Data Collection 📊 Instead of using teleoperation or scripted policies, VRB can allow for an automatic collection of good quality interaction-rich data for policy learning.

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

2. Bootstrapping Exploration 🔍 Robots can utilize the predicted affordances from VRB to explore their environment more intelligently, leading to the discovery of novel and effective ways to manipulate objects.

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

3. Goal-Conditioned Learning 🎯 Can we use our affordance model to iteratively improve performance? We train goal-oriented policies using actions sampled from VRB. We prune this action distribution using the distance to the goal during the training process (using WHIRL RSS'22).

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

4. Action Space Reparameterization 🔄 We can also treat our affordances as an action space itself. It helps accelerate online RL eliminating the need for task-specific primitives. This enables DQN to learn manipulation tasks in under 30 mins!

Deepak Pathak profil fotoğrafı
Deepak Pathak3 yıl önce

Using internet-scale videos is a promising approach to tackling the data problem in robotics and we hope VRB is a strong step toward that! Led by @shikharbahl, @mendonca_rl, @lchen915, @unnatjain2010 CVPR 2023 Paper: Website: 9/9

Murtaza Dalal profil fotoğrafı
Murtaza Dalal3 yıl önce

Awesome work by my colleagues @shikharbahl, @mendonca_rl, @lchen915! Extracting affordances from human video enables robots to efficiently solve a wide array of manipulation tasks in the real world 🤖. Excited to see where this goes next!

Siyuan HUANG profil fotoğrafı
Siyuan HUANG3 yıl önce

Great ideas and a very interesting presentation. :-) I have checked your website but find the dataset is coming soon, would that be available soon?

Benzer Videolar

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,576 görüntüleme • 1 ay önce