Loading video...

Video Failed to Load

Go Home

We present HDMI, a simple and general framework for learning whole-body interaction skills directly from human videos — no manual reward engineering, no task-specific pipelines. 🤖 67 door traversals, 6 real-world tasks, 14 in simulation. 🔗

135,400 views • 1 year ago •via X (Twitter)

31 Comments

Haoyang Weng's profile picture
Haoyang Weng1 year ago

How it works: 1️⃣ Extract human & object motion from monocular RGB videos 2️⃣ Train RL policies with: • unified object representation • residual action space • interaction reward 3️⃣ Deploy zero-shot to real humanoids

Haoyang Weng's profile picture
Haoyang Weng1 year ago

HDMI is capable of: ✅ 67 consecutive door runs (34 minutes!) ✅ Box loco-manipulation with whole-body coordination ✅ Complex multi-stage behaviors like Truman’s Bow ⚡... all with a single set of rewards and observations!

PrismaX's profile picture
PrismaX1 year ago

Why is it wearing boxing gloves?

Haoyang Weng's profile picture
Haoyang Weng1 year ago

just for protecting the wrist motors. the hardware can be damaged if the robot falls.

R2rule's profile picture
R2rule1 year ago

Nice. But why the trainer human needs to move like a robot? shouldn't it be the other way around?

Bill Chambers's profile picture
Bill Chambers1 year ago

why does the robot have kneepads and ankle socks on?

Lyceum's profile picture
Lyceum10 months ago

Wild how far video-based learning has come, how robust it is to messy real-world footage and occlusions.

Chris Paxton's profile picture
Chris Paxton1 year ago

Cool stuff

Allison Smith's profile picture
Allison Smith14 days ago

There's already something called HDMI. Pick a different name.

Chaim Itself's profile picture
Chaim Itself1 year ago

HDMI? Surely you could have chosen an acronym that doesn't conflict with an existing standard name

Paleo-Reactionary Groyper's profile picture
Paleo-Reactionary Groyper1 year ago

@Indian_Bronson Hey that name is already taken

Christopher Cook's profile picture
Christopher Cook1 year ago

Can you change the name to something unique, surely it deserves that.

Yitang Li's profile picture
Yitang Li1 year ago

Congrats Haoyang!

Haoyang Weng's profile picture
Haoyang Weng1 year ago

Also congrats to you!

Xiatao Sun's profile picture
Xiatao Sun1 year ago

Impressive work! Learning whole-body skills directly from monocular videos without manual reward engineering is exactly what the field needs.

Carlos DP 🤖🇺🇸's profile picture
Carlos DP 🤖🇺🇸1 year ago

Very cool work!

Haoyang Weng's profile picture
Haoyang Weng1 year ago

@the_carlosdp thank you so much!

CIX 🦾's profile picture
CIX 🦾1 year ago

great work man

SoloTech's profile picture
SoloTech1 year ago

This is a big step toward generalizable robotics.

Wealth Archives's profile picture
Wealth Archives1 year ago

@REALCULTNEWS china ai is better rn example

PeterSullivanish's profile picture
PeterSullivanish1 year ago

Great work! 👏👏🎉

Samvel's profile picture
Samvel1 year ago

Gumarth

Ruslan Sergeev's profile picture
Ruslan Sergeev11 months ago

@ElijahGalahad Huge congratulations on the publication! It really is a great publication! Could you please explain the purpose of the Teacher and the Student policies in the repository? These are not quite straightforward and not described in the paper. Thanks!

Haoyang Weng's profile picture
Haoyang Weng11 months ago

Thanks for your interest! ROA stands for regularized online adaptation from arxiv 2210.10044. It's for using more information during teacher training and distill them for student. It may speedup the entire training and help exploration, though direct train could also work.

Ruslan Sergeev's profile picture
Ruslan Sergeev11 months ago

Thank you very much for the swift reply!

Samuel Friday's profile picture
Samuel Friday1 year ago

I think edge cases are the downside But this kind of data could also be mass collected and used for training

Oli's profile picture
Oli1 year ago

this is incredibly impressive the walking is perfect its really smooth it can use stairs move various objects this is a massive step for general robotics

Jude Onyenze's profile picture
Jude Onyenze1 year ago

YOOOO, 3rd Person Video for Whole Body Humanoid Control!! Super cool, were you able to deploy for all tasks?

Haoyang Weng's profile picture
Haoyang Weng1 year ago

I deployed 6 out of 14 tasks. mostly because of hardware limits, e.g. we cannot put mocap markers on a ball/we do not have a foldchair that have its back fixed.

Jude Onyenze's profile picture
Jude Onyenze1 year ago

It still amazing results, maybe the hardware limits can be fixed with new models

MyDick's profile picture
MyDick1 year ago

Pickup stuff carry wounded and disabled people, rescue trapped animals, do a personal security to grab weapons on trains and shield the innocent. That should be the main focus not that karate bullshit the other company was promoting

Related Videos