Loading video...
Video Failed to Load
How can we move beyond static-arm lab setups and learn robot policies in our messy homes? We introduce HoMeR, an imitation learning agent for in-the-wild mobile manipulation. 🧵1/8
42,482 views • 1 year ago •via X (Twitter)
12 Comments

In real homes, both data collection and policy learning are hard — workspaces are large, task horizons are long, and scenes are too diverse to fully cover. Many tasks, like wiping a table or opening a door, also require coordinated whole-body movement, not just arm motion. 🧵2/8

We address these challenges with HoMeR, a hybrid IL agent that leverages: • point clouds + images as input And outputs • absolute poses for long-range movement • relative poses for fine-grained manip all executed via a whole-body controller. 🧵3/8

Specifically, HoMeR builds on our prior work SPHINX: extending it to @jimmyyhwu‘s Tidybot++: with a whole-body controller based on @kevin_zakka‘s mink! 🧵4/8

Across 6 simulated and real-world tasks, HoMeR achieves 79% success, outperforming baselines that lack hybrid actions or whole-body control by 29% 🧵5/8

HoMER also can be conditioned on keypoints derived from VLMs, enabling generalization to unseen visual appearances and clutter. 🧵6/8

We also release an iphone-based interface for intuitive whole-body teleoperation, with self-collision avoidance between the arm, base, and external camera mounts. 🧵7/8

Most fun I’ve had working from home :P Big thanks to the entire team! Rhea Malhotra, Phillip Miao, @yjy0625 , @jimmyyhwu , @HengyuanH, @contactrika, @FrancisEngelman, @DorsaSadigh, @leto__jean paper📄: website + code🔗: 🧵8/8

Which Machine Learning model can beat the market? Check out this post on my free Substack where I share code and commentary for several strategies that beat passively holding a tech stock.

Really cool work!

Thank you!!

This is awesome! Congrats Priya!

Thanks Neil :))
