Загрузка видео...
Не удалось загрузить видео
We present HDMI, a simple and general framework for learning whole-body interaction skills directly from human videos — no manual reward engineering, no task-specific pipelines. 🤖 67 door traversals, 6 real-world tasks, 14 in simulation. 🔗
135,400 просмотров • 1 год назад •via X (Twitter)
Комментарии: 31

How it works: 1️⃣ Extract human & object motion from monocular RGB videos 2️⃣ Train RL policies with: • unified object representation • residual action space • interaction reward 3️⃣ Deploy zero-shot to real humanoids

HDMI is capable of: ✅ 67 consecutive door runs (34 minutes!) ✅ Box loco-manipulation with whole-body coordination ✅ Complex multi-stage behaviors like Truman’s Bow ⚡... all with a single set of rewards and observations!

Why is it wearing boxing gloves?

just for protecting the wrist motors. the hardware can be damaged if the robot falls.

Nice. But why the trainer human needs to move like a robot? shouldn't it be the other way around?

why does the robot have kneepads and ankle socks on?

Wild how far video-based learning has come, how robust it is to messy real-world footage and occlusions.

Cool stuff

There's already something called HDMI. Pick a different name.

HDMI? Surely you could have chosen an acronym that doesn't conflict with an existing standard name

@Indian_Bronson Hey that name is already taken

Can you change the name to something unique, surely it deserves that.

Congrats Haoyang!

Also congrats to you!

Impressive work! Learning whole-body skills directly from monocular videos without manual reward engineering is exactly what the field needs.

Very cool work!

@the_carlosdp thank you so much!

great work man

This is a big step toward generalizable robotics.

@REALCULTNEWS china ai is better rn example

Great work! 👏👏🎉

Gumarth

@ElijahGalahad Huge congratulations on the publication! It really is a great publication! Could you please explain the purpose of the Teacher and the Student policies in the repository? These are not quite straightforward and not described in the paper. Thanks!

Thanks for your interest! ROA stands for regularized online adaptation from arxiv 2210.10044. It's for using more information during teacher training and distill them for student. It may speedup the entire training and help exploration, though direct train could also work.

Thank you very much for the swift reply!

I think edge cases are the downside But this kind of data could also be mass collected and used for training

this is incredibly impressive the walking is perfect its really smooth it can use stairs move various objects this is a massive step for general robotics

YOOOO, 3rd Person Video for Whole Body Humanoid Control!! Super cool, were you able to deploy for all tasks?

I deployed 6 out of 14 tasks. mostly because of hardware limits, e.g. we cannot put mocap markers on a ball/we do not have a foldchair that have its back fixed.

It still amazing results, maybe the hardware limits can be fixed with new models

Pickup stuff carry wounded and disabled people, rescue trapped animals, do a personal security to grab weapons on trains and shield the innocent. That should be the main focus not that karate bullshit the other company was promoting
