Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We present HDMI, a simple and general framework for learning whole-body interaction skills directly from human videos — no manual reward engineering, no task-specific pipelines. 🤖 67 door traversals, 6 real-world tasks, 14 in simulation. 🔗

135,400 Aufrufe • vor 1 Jahr •via X (Twitter)

31 Kommentare

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

How it works: 1️⃣ Extract human & object motion from monocular RGB videos 2️⃣ Train RL policies with: • unified object representation • residual action space • interaction reward 3️⃣ Deploy zero-shot to real humanoids

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

HDMI is capable of: ✅ 67 consecutive door runs (34 minutes!) ✅ Box loco-manipulation with whole-body coordination ✅ Complex multi-stage behaviors like Truman’s Bow ⚡... all with a single set of rewards and observations!

Profilbild von PrismaX
PrismaXvor 1 Jahr

Why is it wearing boxing gloves?

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

just for protecting the wrist motors. the hardware can be damaged if the robot falls.

Profilbild von R2rule
R2rulevor 1 Jahr

Nice. But why the trainer human needs to move like a robot? shouldn't it be the other way around?

Profilbild von Bill Chambers
Bill Chambersvor 1 Jahr

why does the robot have kneepads and ankle socks on?

Profilbild von Lyceum
Lyceumvor 10 Monaten

Wild how far video-based learning has come, how robust it is to messy real-world footage and occlusions.

Profilbild von Chris Paxton
Chris Paxtonvor 1 Jahr

Cool stuff

Profilbild von Allison Smith
Allison Smithvor 14 Tagen

There's already something called HDMI. Pick a different name.

Profilbild von Chaim Itself
Chaim Itselfvor 1 Jahr

HDMI? Surely you could have chosen an acronym that doesn't conflict with an existing standard name

Profilbild von Paleo-Reactionary Groyper
Paleo-Reactionary Groypervor 1 Jahr

@Indian_Bronson Hey that name is already taken

Profilbild von Christopher Cook
Christopher Cookvor 1 Jahr

Can you change the name to something unique, surely it deserves that.

Profilbild von Yitang Li
Yitang Livor 1 Jahr

Congrats Haoyang!

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

Also congrats to you!

Profilbild von Xiatao Sun
Xiatao Sunvor 1 Jahr

Impressive work! Learning whole-body skills directly from monocular videos without manual reward engineering is exactly what the field needs.

Profilbild von Carlos DP 🤖🇺🇸
Carlos DP 🤖🇺🇸vor 1 Jahr

Very cool work!

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

@the_carlosdp thank you so much!

Profilbild von CIX 🦾
CIX 🦾vor 1 Jahr

great work man

Profilbild von SoloTech
SoloTechvor 1 Jahr

This is a big step toward generalizable robotics.

Profilbild von Wealth Archives
Wealth Archivesvor 1 Jahr

@REALCULTNEWS china ai is better rn example

Profilbild von PeterSullivanish
PeterSullivanishvor 1 Jahr

Great work! 👏👏🎉

Profilbild von Samvel
Samvelvor 1 Jahr

Gumarth

Profilbild von Ruslan Sergeev
Ruslan Sergeevvor 11 Monaten

@ElijahGalahad Huge congratulations on the publication! It really is a great publication! Could you please explain the purpose of the Teacher and the Student policies in the repository? These are not quite straightforward and not described in the paper. Thanks!

Profilbild von Haoyang Weng
Haoyang Wengvor 11 Monaten

Thanks for your interest! ROA stands for regularized online adaptation from arxiv 2210.10044. It's for using more information during teacher training and distill them for student. It may speedup the entire training and help exploration, though direct train could also work.

Profilbild von Ruslan Sergeev
Ruslan Sergeevvor 11 Monaten

Thank you very much for the swift reply!

Profilbild von Samuel Friday
Samuel Fridayvor 1 Jahr

I think edge cases are the downside But this kind of data could also be mass collected and used for training

Profilbild von Oli
Olivor 1 Jahr

this is incredibly impressive the walking is perfect its really smooth it can use stairs move various objects this is a massive step for general robotics

Profilbild von Jude Onyenze
Jude Onyenzevor 1 Jahr

YOOOO, 3rd Person Video for Whole Body Humanoid Control!! Super cool, were you able to deploy for all tasks?

Profilbild von Haoyang Weng
Haoyang Wengvor 1 Jahr

I deployed 6 out of 14 tasks. mostly because of hardware limits, e.g. we cannot put mocap markers on a ball/we do not have a foldchair that have its back fixed.

Profilbild von Jude Onyenze
Jude Onyenzevor 1 Jahr

It still amazing results, maybe the hardware limits can be fixed with new models

Profilbild von MyDick
MyDickvor 1 Jahr

Pickup stuff carry wounded and disabled people, rescue trapped animals, do a personal security to grab weapons on trains and shield the innocent. That should be the main focus not that karate bullshit the other company was promoting

Ähnliche Videos