Загрузка видео...

Не удалось загрузить видео

На главную

1/ 🧠Humans are the best robot data source! 2/ 👓Human egocentric video is rich in quantity, but poor in quality. 3/ Beyond scaling data, smarter representation and architecture matter just as much. 4/ Want an open-source framework to train your own learn-from-human-data robot policy? 🚀We introduce HumanEgo: Zero-Shot Robot...

114,334 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 56

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「One Policy, Many Conditions」 1️⃣ Cross-Embodiment 2️⃣Cross-Environment 3️⃣Cross-Setup HumanEgo maintains robust success across different conditions without retraining. 🧵 2/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🔥HOT TAKE - 1/3」 🔥Explicit spatial representation, NOT visual fidelity, is the key to bridging the embodiment gap. 🔥Neither hand nor object alone defines a skill—what matters is their interaction. So we design Interaction-Centric Tokens (ICT), a compact entity-level representation of hand–object inter-action invariant to embodiment, viewpoint, and environment. 🤔Most human manipulation is NOT really vision-driven — try putting a flower into a vase with your eyes closed. We do it on spatial memory, proprioception, touch, and hand–object geometry. ICT is built around exactly those signals — which we believe is why HumanEgo generalizes so widely. 🧵 3/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🔥HOT TAKE - 2/3」 🔥Beyond action labels, human videos contain way more information. Richer supervision signal provides complementary gains. 🔥So we add three auxiliary objectives — all sharing the policy's encoder: 1️⃣Object Motion - forecast each object's 6-DoF future) 2️⃣2D Trace - forecast hand/object image-plane paths 3️⃣Latent Consistency - forecast the encoder's own state K steps ahead Together they turn the encoder into a lightweight world model of hand–object interaction — letting us squeeze every demo for every bit of signal, with NO extra labels, NO extra data, just more questions asked of the same trajectory. 🧵 4/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🔥HOT TAKE - 3/3」 🔥Human video is NOT merely a cheap substitute but a superior and more efficient data source for policy learning. Two reasons human data wins: 1️⃣ Higher offline data quality — humans naturally produce smoother, faster, more dexterous motion across a much larger workspace, switching grip strategies in milliseconds with no control loop. Teleop bottlenecks on the operator's piloting skill — everyday teleop is choppy, slow, unrepresentative. 2️⃣ Higher data efficiency — with the right representation (ICT), one human dataset transfers across embodiments, cameras, and environments. Teleop data is born inside one robot body — switch arm, gripper, or camera mount, re-collect everything. 🔥Smoother, broader, transferable — that's not a substitute, that's an upgrade. 🧵 5/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「👓Data Collection & Preprocessing with Aria Glasses @meta_aria 」 ✦ Collecting — Anyone, anytime, anywhere — wear Aria glasses, do the task naturally. No calibration, no special workspace, no lighting setup. 30 min per task. ✦ Preprocessing — Aria's MPS pipeline gives us 6-DoF SLAM + 3D hand pose + synchronized egocentric RGB, all from one lightweight wearable. On top, we use off-the-shelf foundation models for lightweight, accurate object tracking — completing the hand–object pair ICT needs. Why Aria — nothing else in the category checks both boxes today: 1️⃣ Lightweight, production-grade hardware — wearable naturally for long sessions, in any environment 2️⃣ Turnkey MPS pipeline — calibrated SLAM, hand pose, multi-stream RGB, no per-session calibration 3️⃣Accuracy isn't a luxury — our experiments show upstream pose accuracy directly bounds downstream policy performance. Aria's clean, drift-free trajectories are exactly what let HumanEgo converge on minutes of data. 🧵 6/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🗺️System Overview」 🔥Arm inpainting and keypoints rendering bridge the visual gap. 🔥Interaction-Centric Tokens (ICT) encode spatial relationships among all entities 🔥A flow matching policy with dense auxiliary objectives learns bimanual robot actions from minutes-scale human data. 🧵 7/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🤖HumanEgo Bridges the Embodiment Gap Efficiently」 Real-world success rate (%) for each method across all four tasks with 30 min of data. 「HumanEgo / ACT / SPOT / ZeroMimic / Tract2Act / PointPolicy / EgoZero」Three results: 1️⃣HumanEgo achieves the highest success rate on every single task. 2️⃣Even with half the data, HumanEgo outperforms robot teleoperation. 3️⃣HumanEgo excels on tasks that demand precise coordination and spatial reasoning. 🧵 8/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🤯More Tasks」 🔥Real talk — the first time the robot nailed these tasks from just 30 min of casual human-egocentric data (no fine-tuning, no robot data, no per-task tweaking), I was honestly a little stunned. It just worked. 🔥And the model is light enough that we routinely train 10 jobs in parallel on a single RTX 4090. One commodity GPU. Ten experiments at once. That's how hard the system squeezes signal out of limited data. 🧵 9/n

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

「🙏Thank you! 」 Huge shoutout to my amazing collaborators @BotaoUMD, @ColinYu14116982 , @JayLEE_0301 Deepest thanks to my advisors @RuohanGao1 , @furongh , @YAloimonos for the guidance and support throughout. And to everyone in @prgumd and Furong's lab at @umdcs — thanks for the help. 🌐 Website: 📄 Paper: 💻 Code: 📹 Video: 🧵 10/n

Фото профиля Omar Rayyan
Omar Rayyan4 месяцев назад

Congrats on the release. quick question, did you ablate feeding the policy the spatial observations (ict tokens) only without any rgb? judging by the "human rgb" vs "human rgb + ict" gap it seems like that's all what's being used?

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Hi Omar, great question! Yes, we did run this ablation — ICT-only (no RGB) actually performs quite well, around 70%. The intuition is that ICT reframes the problem from "inferring 3D dynamics from pixels" to "learning actions from explicit spatial relationships", so most of the manipulation signal is already encoded in the tokens. It's kind of like glancing at a scene, closing your eyes, and still being able to place bread on a plate — the spatial layout is enough to get you most of the way there. That said, the RGB still pulls its weight — it contributes to fine-grained precision and real-time feedback during contact, which is why adding it nudges performance above the ICT-only baseline. Great question!

Фото профиля Emerson S
Emerson S4 месяцев назад

Every week robotics gets more insane

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

LOL true words

Фото профиля Chris Liu
Chris Liu4 месяцев назад

This is awesome, congrats

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Chris!

Фото профиля Sudhir Pratap Yadav
Sudhir Pratap Yadav4 месяцев назад

Hi I read paper great work. I had one question did you compared Robot teleop data performance but with X axis as data time (~num trajectories) instead of collection time. So for example the figure 5 in paper but X axis is num trajectories And same with ACT

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Hi Sudhir, thanks for reading! Great question. Per trajectory, teleop and human collection time is roughly a 3:2 ratio — e.g., for Serve Bread, one teleop demo takes ~60s while one human demo takes <40s. So you could indeed re-plot Fig. 5 with x-axis = number of trajectories using this conversion. But the takeaway doesn't change: even at the same trajectory count, the human-data policy still outperforms across the board. A few reasons: Human motions are cleaner and smoother, which makes them easier to learn from and transfer. Teleop policies (ACT) mostly rely on the wrist cameras rather than the head-mounted one, whereas HumanEgo only has access to the egocentric (head) view — so the comparison setups are also quite different. Hope this helps!

Фото профиля Sudhir Pratap Yadav
Sudhir Pratap Yadav4 месяцев назад

Thanks for explanation, yes it answer it.

Фото профиля Oleg
Oleg4 месяцев назад

Great post and a very interesting approach. Once I discovered Meta Aria, all bets went for it.

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you! Aria is definitely the best data collection hardware currently.

Фото профиля Oleg
Oleg4 месяцев назад

Would love to collect data using Aria. We applied to program - but no response yet. Have big pool of trained operators to do this

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

🥲True. There is indeed a great demand for the glasses. And Aria Gen2 comes out. It’s probably easier to apply for.

Фото профиля Shenyuan Gao
Shenyuan Gao4 месяцев назад

Congrats for the great work!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Shenyuan!

Фото профиля John Dagdelen
John Dagdelen2 месяцев назад

I noticed that the human data collector moves their hands very slowly in the training examples that you guys released. Way slower that the speed people normally would complete tasks with. Did you try this with faster movements? Is there a reason it was collected that way?

Фото профиля Caesar
Caesar4 месяцев назад

a impressive job!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you!

Фото профиля Steven Cheng
Steven Cheng4 месяцев назад

This data efficiency is huge for my delta arm builds. Just thirty minutes of egocentric video and zero robot data really streamlines the whole training loop.

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Hopefully it will be useful for your project!

Фото профиля Ashok
Ashok4 месяцев назад

This is soo cool!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you!

Фото профиля Jiafei Duan
Jiafei Duan3 месяцев назад

Is the pre-requisite to use this is to have an aria glasses?

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang3 месяцев назад

Yes, it works on only aria glasses for now. We're also trying to support more platform like iPhone or Realsense.

Фото профиля Jingrui Pang
Jingrui Pang4 месяцев назад

cool!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Jingrui!

Фото профиля Winston Gao
Winston Gao4 месяцев назад

Nice work, Leo!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Jiawei!

Фото профиля Chengxuan Qian
Chengxuan Qian4 месяцев назад

Nice work!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Chengxuan!

Фото профиля Bobir M
Bobir M4 месяцев назад

Seems impressive

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you!

Фото профиля Joseph Sottile 🌊 | Diffraction
Joseph Sottile 🌊 | Diffraction4 месяцев назад

This is amazing to see! Data is power 🌊

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Joseph!

Фото профиля Zhang Wang
Zhang Wang4 месяцев назад

Maybe use a dexterous hand

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Yup, that's my next plan!

Фото профиля Alfred Cueva
Alfred Cueva4 месяцев назад

Impressive work, Leo! Leveraging priors from egocentric videos and transferring them across embodiments is a very exciting direction.

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Alfred!

Фото профиля Angkul
Angkul4 месяцев назад

Awesome work. Specially the ICT representation is a smart decisions. And it is highly relevant to my work. Will take it for a spin. Thanks for the code.

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you!

Фото профиля Jie Wang
Jie Wang4 месяцев назад

That’s awesome, congrats Leo!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Tony!

Фото профиля AUTONOMOUS
AUTONOMOUS4 месяцев назад

A lot of these "hot takes" are soon to be the accepted industry ideas! Chef Robotics probably knows the skill defiance better than anyone, they built an LLM to build out the data behind the interaction Great research

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Lovely words! Thank you!

Фото профиля Marilyn Liu
Marilyn Liu4 месяцев назад

Congrats!

Фото профиля Zhi (Leo) Wang
Zhi (Leo) Wang4 месяцев назад

Thank you, Marilyn!

Фото профиля Cowboy🔶BNB |版本之牛
Cowboy🔶BNB |版本之牛4 месяцев назад

Please DM me, I'd like to invest in your project.

Похожие видео

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,681 просмотров • 1 месяц назад