Video yükleniyor...
Video Yüklenemedi
🤔 Can we train one policy to control a wide range of robots, from drones to quadrupeds, navigators to bimanual manipulators, and more? 🦾Introducing CrossFormer: a single policy that can perform manipulation, navigation, aviation, and locomotion:
85,220 görüntüleme • 2 yıl önce •via X (Twitter)
13 Yorum

Cross-embodiment policy learning poses many challenges. Robots may have different control control frequencies, observation spaces (proprioception vs images), and action spaces (2D navigation vs 14D bimanual manipulation)

The core of our method is a transformer-based policy. By framing cross-embodiment training as a sequence-to-sequence problem, we’re able to flexibly adapt our model to robots of varying observation and action spaces

Our model’s flexible nature allows us to outperform and at least match state-of-the-art performance across six embodiments, ranging from bimanual manipulators, navigators, quadrupeds, and more!

We also outperform the best prior cross-embodiment learning work by 3x. Prior work aligns observation & action spaces for navigation and manipulation; we hypothesize that this alignment can sometimes hinder performance on more complex tasks

Check out our website and code for more details on CrossFormer! Website: Paper: Code:

Really grateful and lucky to have worked with some incredible people on this project: co-lead @HomerWalke and collaborators @oier_mees @SudeepDasari @svlevine

Cool

The paper aims to train a single policy that can control a diverse set of robot embodiments, including single and bimanual robot arms, ground navigation robots, quadcopters, and quadrupeds. This is a challenging task as robots can have widely varying sensors, actuators, and control frequencies. The researchers find that their CrossFormer policy matches the performance of the same architecture trained only on the target robot's data, as well as the best prior method in each evaluation setting. CrossFormer outperforms the prior state-of-the-art in cross-embodiment learning by a large margin, without requiring manual alignment of observation and action spaces. full paper:

@davide_tateo

Super cool work, congrats Ria!

Very cool.

Absolutely brilliant. Will check this out

Interesting work👍
