正在加载视频...

视频加载失败

R+X was accepted at ICRA 2025! Robots can now do in-context imitation learning, just by observing humans going about their daily lives... No more need to *label* and *train* - just *RETRIEVE* and *EXECUTE*! 🧵👇 (1/5)

10,402 次观看 • 1 年前 •via X (Twitter)

6 条评论

Edward Johns 的头像
Edward Johns1 年前

Given a language command at deployment: (1) Use Gemini to retrieve all the relevant human videos and their human hand trajectories, (2) Condition our in-context IL method (KAT) on these trajectories, (3) Predict the robot hand trajectories, and execute! (2/5)

Edward Johns 的头像
Edward Johns1 年前

Importantly: (1) We don't require the human videos to be labelled. (2) Since we do in-context imitation learning "at test time" rather than training an explicit policy, new human videos can easily be added on the fly and used immediately for retrieval. This is very scalable! (3/5)

Edward Johns 的头像
Edward Johns1 年前

We also found that this "Retrieval + Execution" idea performs much better than training an explicit, language-conditioned policy, such as when fine-tuning R3M or Octo on this same dataset of human videos. (4/5)

Edward Johns 的头像
Edward Johns1 年前

R+X was jointly led by Georgios Papagiannis (@geopgs) and Norman Di Palo (@normandipalo). For the paper and further videos, please visit: Thanks for reading! (5/5)

The Rundown AI 的头像
The Rundown AI2 年前

If you're not learning AI in 2024, you're falling behind. Join 500,000+ readers and learn how to use AI in just 5 minutes a day (for free).

Michael Cho - Rbt/Acc 的头像
Michael Cho - Rbt/Acc1 年前

Super cool work! 2 qn: must the human videos be recorded at the same environment where the physical robots do in-context learning at inference? Also, how important is the Depth info (saw that u guys are using rgbd cameras in this setup)?

相关视频