正在加载视频...

视频加载失败

Imitation learning has a data scarcity problem. Introducing EgoDex from Apple, the largest and most diverse dataset of dexterous human manipulation to date — 829 hours of egocentric video + paired 3D hand poses across 194 tasks. Now on arxiv: (1/4)

114,164 次观看 • 1 年前 •via X (Twitter)

11 条评论

Ryan Hoque 的头像
Ryan Hoque1 年前

Unlike teleoperation, egocentric video is passively scalable - like text and images on the Internet. We use Apple Vision Pro to collect video + precise pose annotations (unlike Ego4D, which lacks native pose data). This unlocks 5x the scale of existing large datasets like DROID.

Ryan Hoque 的头像
Ryan Hoque1 年前

We also propose new benchmarks and train imitation learning policies for dexterous trajectory prediction. Below are 30 Hz wrist and fingertip trajectories on the test set, where blue = ground truth, red = model predictions, and points get lighter up to 2 seconds in the future.

Ryan Hoque 的头像
Ryan Hoque1 年前

The full dataset is now publicly available to the community, access details are in the paper. Sample code for data loading is coming soon. Enjoy!

SecBriefs | Making Cybersecurity Simple 的头像
SecBriefs | Making Cybersecurity Simple1 年前

⚠️The average person generates 2.5 quintillion bytes of data annually. That's enough to fill 575,000 libraries!📚 This data is used to track, target, and manipulate you. #Cybersecurity matters.💡 Cybersecurity Dictionary for Everyone is on Apple Books:

Raul 的头像
Raul1 年前

Perfect for Optimus to learn new skills

RL 的头像
RL1 年前

Fyi the dataset links dont work: “ NoSuchKeyThe specified key does not exist.datasets/egodex/[filename].zip9C2FBJJ7FJHKHDT3Uk3l1oHoR9NeaNJdC7gInDjt5u8slFtW5lRt9wFR0MQIWNXIk4sTWiLGEYF22KUPQQ9X6CVC+UU=”

Michael Black 的头像
Michael Black1 年前

Looks great. You mention that it’s now public but I don’t find the link anywhere.

Hussein Lezzaik 的头像
Hussein Lezzaik1 年前

excellent work, congrats!

Soroush Nasiriany 的头像
Soroush Nasiriany1 年前

Congrats Ryan! Awesome work as always!!

Humanoids daily 的头像
Humanoids daily1 年前

Impressive! Interesting use of the Apple Vision Pro.

Idriel Vermillion 的头像
Idriel Vermillion1 年前

So hype

相关视频

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

293,383 次观看 • 4 个月前