Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We present HDMI, a simple and general framework for learning whole-body interaction skills directly from human videos — no manual reward engineering, no task-specific pipelines. 🤖 67 door traversals, 6 real-world tasks, 14 in simulation. 🔗

135,400 görüntüleme • 1 yıl önce •via X (Twitter)

31 Yorum

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

How it works: 1️⃣ Extract human & object motion from monocular RGB videos 2️⃣ Train RL policies with: • unified object representation • residual action space • interaction reward 3️⃣ Deploy zero-shot to real humanoids

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

HDMI is capable of: ✅ 67 consecutive door runs (34 minutes!) ✅ Box loco-manipulation with whole-body coordination ✅ Complex multi-stage behaviors like Truman’s Bow ⚡... all with a single set of rewards and observations!

PrismaX profil fotoğrafı
PrismaX1 yıl önce

Why is it wearing boxing gloves?

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

just for protecting the wrist motors. the hardware can be damaged if the robot falls.

R2rule profil fotoğrafı
R2rule1 yıl önce

Nice. But why the trainer human needs to move like a robot? shouldn't it be the other way around?

Bill Chambers profil fotoğrafı
Bill Chambers1 yıl önce

why does the robot have kneepads and ankle socks on?

Lyceum profil fotoğrafı
Lyceum10 ay önce

Wild how far video-based learning has come, how robust it is to messy real-world footage and occlusions.

Chris Paxton profil fotoğrafı
Chris Paxton1 yıl önce

Cool stuff

Allison Smith profil fotoğrafı
Allison Smith14 gün önce

There's already something called HDMI. Pick a different name.

Chaim Itself profil fotoğrafı
Chaim Itself1 yıl önce

HDMI? Surely you could have chosen an acronym that doesn't conflict with an existing standard name

Paleo-Reactionary Groyper profil fotoğrafı
Paleo-Reactionary Groyper1 yıl önce

@Indian_Bronson Hey that name is already taken

Christopher Cook profil fotoğrafı
Christopher Cook1 yıl önce

Can you change the name to something unique, surely it deserves that.

Yitang Li profil fotoğrafı
Yitang Li1 yıl önce

Congrats Haoyang!

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

Also congrats to you!

Xiatao Sun profil fotoğrafı
Xiatao Sun1 yıl önce

Impressive work! Learning whole-body skills directly from monocular videos without manual reward engineering is exactly what the field needs.

Carlos DP 🤖🇺🇸 profil fotoğrafı
Carlos DP 🤖🇺🇸1 yıl önce

Very cool work!

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

@the_carlosdp thank you so much!

CIX 🦾 profil fotoğrafı
CIX 🦾1 yıl önce

great work man

SoloTech profil fotoğrafı
SoloTech1 yıl önce

This is a big step toward generalizable robotics.

Wealth Archives profil fotoğrafı
Wealth Archives1 yıl önce

@REALCULTNEWS china ai is better rn example

PeterSullivanish profil fotoğrafı
PeterSullivanish1 yıl önce

Great work! 👏👏🎉

Samvel profil fotoğrafı
Samvel1 yıl önce

Gumarth

Ruslan Sergeev profil fotoğrafı
Ruslan Sergeev11 ay önce

@ElijahGalahad Huge congratulations on the publication! It really is a great publication! Could you please explain the purpose of the Teacher and the Student policies in the repository? These are not quite straightforward and not described in the paper. Thanks!

Haoyang Weng profil fotoğrafı
Haoyang Weng11 ay önce

Thanks for your interest! ROA stands for regularized online adaptation from arxiv 2210.10044. It's for using more information during teacher training and distill them for student. It may speedup the entire training and help exploration, though direct train could also work.

Ruslan Sergeev profil fotoğrafı
Ruslan Sergeev11 ay önce

Thank you very much for the swift reply!

Samuel Friday profil fotoğrafı
Samuel Friday1 yıl önce

I think edge cases are the downside But this kind of data could also be mass collected and used for training

Oli profil fotoğrafı
Oli1 yıl önce

this is incredibly impressive the walking is perfect its really smooth it can use stairs move various objects this is a massive step for general robotics

Jude Onyenze profil fotoğrafı
Jude Onyenze1 yıl önce

YOOOO, 3rd Person Video for Whole Body Humanoid Control!! Super cool, were you able to deploy for all tasks?

Haoyang Weng profil fotoğrafı
Haoyang Weng1 yıl önce

I deployed 6 out of 14 tasks. mostly because of hardware limits, e.g. we cannot put mocap markers on a ball/we do not have a foldchair that have its back fixed.

Jude Onyenze profil fotoğrafı
Jude Onyenze1 yıl önce

It still amazing results, maybe the hardware limits can be fixed with new models

MyDick profil fotoğrafı
MyDick1 yıl önce

Pickup stuff carry wounded and disabled people, rescue trapped animals, do a personal security to grab weapons on trains and shield the innocent. That should be the main focus not that karate bullshit the other company was promoting

Benzer Videolar