正在加载视频...

视频加载失败

Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.

301,754 次观看 • 1 个月前 •via X (Twitter)

32 条评论

Generalist 的头像
Generalist1 个月前

Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical commonsense that transfers to new hands and new ways to grasp, push, pull, twist, and more.

Generalist 的头像
Generalist1 个月前

Every hand is a different language for physical interaction. Just as training on multiple languages produces more capable LLMs, a model trained across embodiments can gather shared knowledge across instances and more readily separate what is specific to a hand from what is universal about the world.

Generalist 的头像
Generalist1 个月前

Unlike humans, robots aren’t locked into the hands they’re born with. Machines can easily swap mechanical hands, and tool changers are common in automation. Switching between different end effectors to reach a goal can be viewed as a form of physical reasoning: using the right tool for the right job, much as multilingual chain-of-thought can improve LLM reasoning for downstream RL.

Generalist 的头像
Generalist1 个月前

What happens when you change the hand mid-task? We tested this by modifying the hands mid-rollout and letting the same model keep running. It perceives the new tool, and finds a new trajectory and contact strategy to complete the task. This works because training on mixed data forces the model to condition its behavior on the hand in front of it. This brings it closer to a general understanding of how shapes and contact surfaces interact with the physical world – and which actuation strategy should follow.

Generalist 的头像
Generalist1 个月前

Nature didn’t converge on just a single solution for manipulating the physical world. It exploded into millions. You can see this everywhere in the diversity of life around us: from the beak of a bird, to the trunk of an elephant, from the suction pads of an octopus, to the pollen baskets of a honeybee. The human hand is only a single point in a design space so large we’ve barely begun to map. For machines, five-fingered hands will just be one tool among many; limiting robots to only that would be a failure of imagination.

Generalist 的头像
Generalist1 个月前

If we can build general intelligence that understands the underlying physics of interaction, then the shape of the hand becomes secondary to the intelligence that drives it. A suction pad, a gripper, a brush, a plasma welding nozzle – are all just different interfaces through which the same intelligence can reshape the physical world. The future of robot hands won’t look like ours. It will look more like a toolbox with a thousand hands: augmented, recombined, and scaled. Robots were always meant to extend what humans can do – to empower people to shape the physical world in places and at scales we never could before. Read more about our work at:

Xianyi Cheng 的头像
Xianyi Cheng1 个月前

Great work! We have a low-cost, 3D-printed quick-swapping system build on a 1 DoF gripper that generalizes across everyday tools We can even attach it to a $7 hand :)

Lane Burgett 的头像
Lane Burgett1 个月前

What about my excavator embodiment?

maru 的头像
maru1 个月前

Gen-1 is Trillion time's better then figure toys

Mikhail 的头像
Mikhail1 个月前

Amazing! We have our own quick change mechanism on our custom build bi-manual robot

John Dagdelen 的头像
John Dagdelen1 个月前

Really cool stuff! One question, though. How does the model adjust the action space it outputs in real time when end effector changes? Does the system have to communicate to the model that an axis changed? For example, when gripper gets taped up.

Clawbrowser - Anti-Detect Browser for AI Agents 的头像
Clawbrowser - Anti-Detect Browser for AI Agents1 个月前

finally an intern who can use both the browser and his hands

Ella Tech & Tool 的头像
Ella Tech & Tool1 个月前

Smart move One brain, any hand. That's how you scale real robot dexterity.

Zane 的头像
Zane1 个月前

More hands, more learning opportunities

Bblack 的头像
Bblack1 个月前

One model, infinite hands 🤖 GEN-1 learning from all end effectors is huge

Akriti 的头像
Akriti1 个月前

This is the future

Emma 的头像
Emma1 个月前

If one policy can absorb a hand swap mid-rollout, embodiment starts looking like context rather than identity. The next leap is letting the robot choose and swap tools itself—turning hardware selection into inference-time reasoning.

SEAR 的头像
SEAR1 个月前

single hand models are cool

Context 的头像
Context1 个月前

The interesting metric isn’t how many end effectors it supports. It’s how much performance survives when you stop optimizing for just one.

Mi.lu. 的头像
Mi.lu.1 个月前

OpenAI announced Project Camellia, a 3.2 GW data-center campus in Georgia. Construction is planned in phases from 2028 to 2032. OpenAI says it will cover the infrastructure costs, use closed-loop water systems, and fund $80M in local benefits. Those are firm commitments. Execution and local impact still need to be demonstrated. Source:

Lex 的头像
Lex1 个月前

The many hands framing makes a lot of sense. Robots probably shouldn’t be forced into one human like hand design when different tasks need different tools.

Guillermo Duró 的头像
Guillermo Duró1 个月前

This is awesome work

Nihal 的头像
Nihal1 个月前

@huaijiangzhu So sick

Andy Wojcicki 的头像
Andy Wojcicki1 个月前

that switch at ~1:00

ArtuR2 ⏩️⤴️ 的头像
ArtuR2 ⏩️⤴️1 个月前

@elonmusk buy them! 🙏🏼

Bela Becerra 的头像
Bela Becerra1 个月前

incredible !!!

Sam W 的头像
Sam W1 个月前

This is incredible. So I can apply this to robots I build? I am working on a cooking robot right now

Vivek Gopalan 的头像
Vivek Gopalan1 个月前

@peteflorence akin to magic

Ramzi RAMZI 的头像
Ramzi RAMZI1 个月前

Can startups use your models on there robots ??

Han 的头像
Han1 个月前

Congratulations!

Julian Fried 的头像
Julian Fried1 个月前

Very cool

Noor Tech 的头像
Noor Tech1 个月前

GEN-1 learns from many, works with all

相关视频

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

300,427 次观看 • 6 个月前