Video wird geladen...
Video konnte nicht geladen werden
For large-scale robotic deployment🤖 in the real-world 🌏, robots must adapt to changes in environment and objects. Ever questioned the generalizability of your robot's manipulation policy? Put it to the test with The Colosseum 🏛️. Check out our project:
36,617 Aufrufe • vor 2 Jahren •via X (Twitter)
14 Kommentare

1/12 🧵: Introducing THE COLOSSEUM: A benchmark for evaluating robot manipulation against environmental changes. With perturbations in 20 tasks from RLBench (@stepjamUK), it spans 12 dimensions of challenges

2/12 🧵: To set the foundation for this new robotics benchmark, we assessed 4 state-of-the-art (SotA) behavior cloning models - R3M, MVP, PerAct @mohito1905 , and RVT @imankitgoyal - across 12 perturbation factors and 20 robotics manipulation tasks. 📊

3/12 🧵: Exploring BC models through THE COLOSSEUM, we delve into key research questions on generalization for BC policies. Are 2D-based BC models more generalizable than 3D? What factors most disrupt BC models' performance? Does the real world align with simulations in terms of model generalization? 🤔

4/12 🧵: Fresh insights from our evaluation results! 🔥We've found that 3D baselines generally outperform 2D ones, showing significantly greater robustness to environmental perturbations.

5/12 🧵: Zeroing in on 3D models (RVT and PerAct), we've identified that color nuances —from objects, tables, to lighting—and distractors 🚨 play pivotal roles in performance. In contrast, 2D models (R3M-MLP and MVP-MLP) show heightened sensitivity towards object and light color, texture, and the angle of the camera 📷.

6/12 🧵: We've brought our simulations to life! We're testing real-world perturbations on a trained BC model by replicating 4 simulation tasks and 3D printing all assets. This enhances our benchmark's reproducibility and bridges the gap between simulation and real-world tasks.

7/12 🧵: Real-world tests align with The Colosseum's simulation evaluation results, showing you can trust BC model sims to reflect similar performance drops under the same perturbation factors⚖️.

8/12 🧵: Also, check out these cool prior works: Decomposing visual imitation learning– – great to have independent works converge on similar ideas and initiatives!

9/12 🧵: The Colosseum will continue to serve as a challenge and provide a unified platform to develop, evaluate, and compare future robotic manipulation methods that stand the test of robustness and generalization.

10/12 🧵: We're dedicated to keeping The Colosseum's leaderboard up-to-date ✨. We will add in more model evals (ACT3D, GNFactor, ACT, and many more). Excited to see new robotic manipulation policies step up and test their generalizability with us!

11/12 🧵: This project could not have been possible without the team @wpumacay7567 @Ishika_S_ , and advisors @RanjayKrishna @_jessethomason_ Dieter Fox.

12/12🧵: Stay tuned for the release of our code and complete benchmark. However, if you would like to evaluate your BC model with us before our release, feel free to contact us at: [email protected]

Exciting benchmarks like ImageNet-A, WILDS, and GLUE have driven advancements in vision and language models by testing for robustness and generalization. It's time for robotics to get a piece of the action and push the field forward with this tailored benchmark!

@allen_ai @nvidia @CSatUSC
