Loading video...

Video Failed to Load

Go Home

Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.

301,754 views • 1 month ago •via X (Twitter)

32 Comments

Generalist's profile picture
Generalist1 month ago

Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical commonsense that transfers to new hands and new ways to grasp, push, pull, twist, and more.

Generalist's profile picture
Generalist1 month ago

Every hand is a different language for physical interaction. Just as training on multiple languages produces more capable LLMs, a model trained across embodiments can gather shared knowledge across instances and more readily separate what is specific to a hand from what is universal about the world.

Generalist's profile picture
Generalist1 month ago

Unlike humans, robots aren’t locked into the hands they’re born with. Machines can easily swap mechanical hands, and tool changers are common in automation. Switching between different end effectors to reach a goal can be viewed as a form of physical reasoning: using the right tool for the right job, much as multilingual chain-of-thought can improve LLM reasoning for downstream RL.

Generalist's profile picture
Generalist1 month ago

What happens when you change the hand mid-task? We tested this by modifying the hands mid-rollout and letting the same model keep running. It perceives the new tool, and finds a new trajectory and contact strategy to complete the task. This works because training on mixed data forces the model to condition its behavior on the hand in front of it. This brings it closer to a general understanding of how shapes and contact surfaces interact with the physical world – and which actuation strategy should follow.

Generalist's profile picture
Generalist1 month ago

Nature didn’t converge on just a single solution for manipulating the physical world. It exploded into millions. You can see this everywhere in the diversity of life around us: from the beak of a bird, to the trunk of an elephant, from the suction pads of an octopus, to the pollen baskets of a honeybee. The human hand is only a single point in a design space so large we’ve barely begun to map. For machines, five-fingered hands will just be one tool among many; limiting robots to only that would be a failure of imagination.

Generalist's profile picture
Generalist1 month ago

If we can build general intelligence that understands the underlying physics of interaction, then the shape of the hand becomes secondary to the intelligence that drives it. A suction pad, a gripper, a brush, a plasma welding nozzle – are all just different interfaces through which the same intelligence can reshape the physical world. The future of robot hands won’t look like ours. It will look more like a toolbox with a thousand hands: augmented, recombined, and scaled. Robots were always meant to extend what humans can do – to empower people to shape the physical world in places and at scales we never could before. Read more about our work at:

Xianyi Cheng's profile picture
Xianyi Cheng1 month ago

Great work! We have a low-cost, 3D-printed quick-swapping system build on a 1 DoF gripper that generalizes across everyday tools We can even attach it to a $7 hand :)

Lane Burgett's profile picture
Lane Burgett1 month ago

What about my excavator embodiment?

maru's profile picture
maru1 month ago

Gen-1 is Trillion time's better then figure toys

Mikhail's profile picture
Mikhail1 month ago

Amazing! We have our own quick change mechanism on our custom build bi-manual robot

John Dagdelen's profile picture
John Dagdelen1 month ago

Really cool stuff! One question, though. How does the model adjust the action space it outputs in real time when end effector changes? Does the system have to communicate to the model that an axis changed? For example, when gripper gets taped up.

Clawbrowser - Anti-Detect Browser for AI Agents's profile picture
Clawbrowser - Anti-Detect Browser for AI Agents1 month ago

finally an intern who can use both the browser and his hands

Ella Tech & Tool's profile picture
Ella Tech & Tool1 month ago

Smart move One brain, any hand. That's how you scale real robot dexterity.

Zane's profile picture
Zane1 month ago

More hands, more learning opportunities

Bblack's profile picture
Bblack1 month ago

One model, infinite hands 🤖 GEN-1 learning from all end effectors is huge

Akriti's profile picture
Akriti1 month ago

This is the future

Emma's profile picture
Emma1 month ago

If one policy can absorb a hand swap mid-rollout, embodiment starts looking like context rather than identity. The next leap is letting the robot choose and swap tools itself—turning hardware selection into inference-time reasoning.

SEAR's profile picture
SEAR1 month ago

single hand models are cool

Context's profile picture
Context1 month ago

The interesting metric isn’t how many end effectors it supports. It’s how much performance survives when you stop optimizing for just one.

Mi.lu.'s profile picture
Mi.lu.1 month ago

OpenAI announced Project Camellia, a 3.2 GW data-center campus in Georgia. Construction is planned in phases from 2028 to 2032. OpenAI says it will cover the infrastructure costs, use closed-loop water systems, and fund $80M in local benefits. Those are firm commitments. Execution and local impact still need to be demonstrated. Source:

Lex's profile picture
Lex1 month ago

The many hands framing makes a lot of sense. Robots probably shouldn’t be forced into one human like hand design when different tasks need different tools.

Guillermo Duró's profile picture
Guillermo Duró1 month ago

This is awesome work

Nihal's profile picture
Nihal1 month ago

@huaijiangzhu So sick

Andy Wojcicki's profile picture
Andy Wojcicki1 month ago

that switch at ~1:00

ArtuR2 ⏩️⤴️'s profile picture
ArtuR2 ⏩️⤴️1 month ago

@elonmusk buy them! 🙏🏼

Bela Becerra's profile picture
Bela Becerra1 month ago

incredible !!!

Sam W's profile picture
Sam W1 month ago

This is incredible. So I can apply this to robots I build? I am working on a cooking robot right now

Vivek Gopalan's profile picture
Vivek Gopalan1 month ago

@peteflorence akin to magic

Ramzi RAMZI's profile picture
Ramzi RAMZI1 month ago

Can startups use your models on there robots ??

Han's profile picture
Han1 month ago

Congratulations!

Julian Fried's profile picture
Julian Fried1 month ago

Very cool

Noor Tech's profile picture
Noor Tech1 month ago

GEN-1 learns from many, works with all

Related Videos

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

300,427 views • 6 months ago