正在加载视频...

视频加载失败

🤖 Robotics often faces a chicken and egg problem: no web-scale robot data for training (unlike CV or NLP) b/c robots aren't deployed yet & vice-versa. Introducing VRB: Use large-scale human videos to train a *general-purpose* affordance model to jumpstart any robotics paradigm!

70,712 次观看 • 3 年前 •via X (Twitter)

10 条评论

Deepak Pathak 的头像
Deepak Pathak3 年前

🌐 Given a new scene, our general-purpose VRB model predicts all the locations *where* a robot can manipulate objects and *how* should it move post-manipulation. 💡Qs: 1) What are affordances? 2) How to get large-scale data? 3) How does this help robots?

Deepak Pathak 的头像
Deepak Pathak3 年前

🧩 In computer vision, there is a long line of work in affordances (Gibson 1966, 1979). However, what's the best way to define them for robotics? Our solution: interaction points & post-contact trajectories, allowing for seamless integration with almost any robot learning setup.

Deepak Pathak 的头像
Deepak Pathak3 年前

🔧🔧 How do we extract affordances from humans? We estimate hand poses & hand-object interaction points to extract contact pts + interaction direction (thanks to @DandanShan_, David Fouhey's 100DoH), and map them back to frames where humans arent present to avoid embodiment gap.

Deepak Pathak 的头像
Deepak Pathak3 年前

How to use this model for different robot learning paradigms? We show 4 different paradigms. 1. Offline Data Collection 📊 Instead of using teleoperation or scripted policies, VRB can allow for an automatic collection of good quality interaction-rich data for policy learning.

Deepak Pathak 的头像
Deepak Pathak3 年前

2. Bootstrapping Exploration 🔍 Robots can utilize the predicted affordances from VRB to explore their environment more intelligently, leading to the discovery of novel and effective ways to manipulate objects.

Deepak Pathak 的头像
Deepak Pathak3 年前

3. Goal-Conditioned Learning 🎯 Can we use our affordance model to iteratively improve performance? We train goal-oriented policies using actions sampled from VRB. We prune this action distribution using the distance to the goal during the training process (using WHIRL RSS'22).

Deepak Pathak 的头像
Deepak Pathak3 年前

4. Action Space Reparameterization 🔄 We can also treat our affordances as an action space itself. It helps accelerate online RL eliminating the need for task-specific primitives. This enables DQN to learn manipulation tasks in under 30 mins!

Deepak Pathak 的头像
Deepak Pathak3 年前

Using internet-scale videos is a promising approach to tackling the data problem in robotics and we hope VRB is a strong step toward that! Led by @shikharbahl, @mendonca_rl, @lchen915, @unnatjain2010 CVPR 2023 Paper: Website: 9/9

Murtaza Dalal 的头像
Murtaza Dalal3 年前

Awesome work by my colleagues @shikharbahl, @mendonca_rl, @lchen915! Extracting affordances from human video enables robots to efficiently solve a wide array of manipulation tasks in the real world 🤖. Excited to see where this goes next!

Siyuan HUANG 的头像
Siyuan HUANG3 年前

Great ideas and a very interesting presentation. :-) I have checked your website but find the dataset is coming soon, would that be available soon?

相关视频

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,576 次观看 • 1 个月前