Loading video...

Video Failed to Load

Go Home

R+X was accepted at ICRA 2025! Robots can now do in-context imitation learning, just by observing humans going about their daily lives... No more need to *label* and *train* - just *RETRIEVE* and *EXECUTE*! 🧵👇 (1/5)

10,402 views • 1 year ago •via X (Twitter)

6 Comments

Edward Johns's profile picture
Edward Johns1 year ago

Given a language command at deployment: (1) Use Gemini to retrieve all the relevant human videos and their human hand trajectories, (2) Condition our in-context IL method (KAT) on these trajectories, (3) Predict the robot hand trajectories, and execute! (2/5)

Edward Johns's profile picture
Edward Johns1 year ago

Importantly: (1) We don't require the human videos to be labelled. (2) Since we do in-context imitation learning "at test time" rather than training an explicit policy, new human videos can easily be added on the fly and used immediately for retrieval. This is very scalable! (3/5)

Edward Johns's profile picture
Edward Johns1 year ago

We also found that this "Retrieval + Execution" idea performs much better than training an explicit, language-conditioned policy, such as when fine-tuning R3M or Octo on this same dataset of human videos. (4/5)

Edward Johns's profile picture
Edward Johns1 year ago

R+X was jointly led by Georgios Papagiannis (@geopgs) and Norman Di Palo (@normandipalo). For the paper and further videos, please visit: Thanks for reading! (5/5)

The Rundown AI's profile picture
The Rundown AI2 years ago

If you're not learning AI in 2024, you're falling behind. Join 500,000+ readers and learn how to use AI in just 5 minutes a day (for free).

Michael Cho - Rbt/Acc's profile picture
Michael Cho - Rbt/Acc1 year ago

Super cool work! 2 qn: must the human videos be recorded at the same environment where the physical robots do in-context learning at inference? Also, how important is the Depth info (saw that u guys are using rgbd cameras in this setup)?

Related Videos