Загрузка видео...

Не удалось загрузить видео

На главную

We present HDMI, a simple and general framework for learning whole-body interaction skills directly from human videos — no manual reward engineering, no task-specific pipelines. 🤖 67 door traversals, 6 real-world tasks, 14 in simulation. 🔗

135,400 просмотров • 1 год назад •via X (Twitter)

Комментарии: 31

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

How it works: 1️⃣ Extract human & object motion from monocular RGB videos 2️⃣ Train RL policies with: • unified object representation • residual action space • interaction reward 3️⃣ Deploy zero-shot to real humanoids

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

HDMI is capable of: ✅ 67 consecutive door runs (34 minutes!) ✅ Box loco-manipulation with whole-body coordination ✅ Complex multi-stage behaviors like Truman’s Bow ⚡... all with a single set of rewards and observations!

Фото профиля PrismaX
PrismaX1 год назад

Why is it wearing boxing gloves?

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

just for protecting the wrist motors. the hardware can be damaged if the robot falls.

Фото профиля R2rule
R2rule1 год назад

Nice. But why the trainer human needs to move like a robot? shouldn't it be the other way around?

Фото профиля Bill Chambers
Bill Chambers1 год назад

why does the robot have kneepads and ankle socks on?

Фото профиля Lyceum
Lyceum10 месяцев назад

Wild how far video-based learning has come, how robust it is to messy real-world footage and occlusions.

Фото профиля Chris Paxton
Chris Paxton1 год назад

Cool stuff

Фото профиля Allison Smith
Allison Smith14 дней назад

There's already something called HDMI. Pick a different name.

Фото профиля Chaim Itself
Chaim Itself1 год назад

HDMI? Surely you could have chosen an acronym that doesn't conflict with an existing standard name

Фото профиля Paleo-Reactionary Groyper
Paleo-Reactionary Groyper1 год назад

@Indian_Bronson Hey that name is already taken

Фото профиля Christopher Cook
Christopher Cook1 год назад

Can you change the name to something unique, surely it deserves that.

Фото профиля Yitang Li
Yitang Li1 год назад

Congrats Haoyang!

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

Also congrats to you!

Фото профиля Xiatao Sun
Xiatao Sun1 год назад

Impressive work! Learning whole-body skills directly from monocular videos without manual reward engineering is exactly what the field needs.

Фото профиля Carlos DP 🤖🇺🇸
Carlos DP 🤖🇺🇸1 год назад

Very cool work!

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

@the_carlosdp thank you so much!

Фото профиля CIX 🦾
CIX 🦾1 год назад

great work man

Фото профиля SoloTech
SoloTech1 год назад

This is a big step toward generalizable robotics.

Фото профиля Wealth Archives
Wealth Archives1 год назад

@REALCULTNEWS china ai is better rn example

Фото профиля PeterSullivanish
PeterSullivanish1 год назад

Great work! 👏👏🎉

Фото профиля Samvel
Samvel1 год назад

Gumarth

Фото профиля Ruslan Sergeev
Ruslan Sergeev11 месяцев назад

@ElijahGalahad Huge congratulations on the publication! It really is a great publication! Could you please explain the purpose of the Teacher and the Student policies in the repository? These are not quite straightforward and not described in the paper. Thanks!

Фото профиля Haoyang Weng
Haoyang Weng11 месяцев назад

Thanks for your interest! ROA stands for regularized online adaptation from arxiv 2210.10044. It's for using more information during teacher training and distill them for student. It may speedup the entire training and help exploration, though direct train could also work.

Фото профиля Ruslan Sergeev
Ruslan Sergeev11 месяцев назад

Thank you very much for the swift reply!

Фото профиля Samuel Friday
Samuel Friday1 год назад

I think edge cases are the downside But this kind of data could also be mass collected and used for training

Фото профиля Oli
Oli1 год назад

this is incredibly impressive the walking is perfect its really smooth it can use stairs move various objects this is a massive step for general robotics

Фото профиля Jude Onyenze
Jude Onyenze1 год назад

YOOOO, 3rd Person Video for Whole Body Humanoid Control!! Super cool, were you able to deploy for all tasks?

Фото профиля Haoyang Weng
Haoyang Weng1 год назад

I deployed 6 out of 14 tasks. mostly because of hardware limits, e.g. we cannot put mocap markers on a ball/we do not have a foldchair that have its back fixed.

Фото профиля Jude Onyenze
Jude Onyenze1 год назад

It still amazing results, maybe the hardware limits can be fixed with new models

Фото профиля MyDick
MyDick1 год назад

Pickup stuff carry wounded and disabled people, rescue trapped animals, do a personal security to grab weapons on trains and shield the innocent. That should be the main focus not that karate bullshit the other company was promoting

Похожие видео