Загрузка видео...
Не удалось загрузить видео
Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
301,754 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 32

Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical commonsense that transfers to new hands and new ways to grasp, push, pull, twist, and more.

Every hand is a different language for physical interaction. Just as training on multiple languages produces more capable LLMs, a model trained across embodiments can gather shared knowledge across instances and more readily separate what is specific to a hand from what is universal about the world.

Unlike humans, robots aren’t locked into the hands they’re born with. Machines can easily swap mechanical hands, and tool changers are common in automation. Switching between different end effectors to reach a goal can be viewed as a form of physical reasoning: using the right tool for the right job, much as multilingual chain-of-thought can improve LLM reasoning for downstream RL.

What happens when you change the hand mid-task? We tested this by modifying the hands mid-rollout and letting the same model keep running. It perceives the new tool, and finds a new trajectory and contact strategy to complete the task. This works because training on mixed data forces the model to condition its behavior on the hand in front of it. This brings it closer to a general understanding of how shapes and contact surfaces interact with the physical world – and which actuation strategy should follow.

Nature didn’t converge on just a single solution for manipulating the physical world. It exploded into millions. You can see this everywhere in the diversity of life around us: from the beak of a bird, to the trunk of an elephant, from the suction pads of an octopus, to the pollen baskets of a honeybee. The human hand is only a single point in a design space so large we’ve barely begun to map. For machines, five-fingered hands will just be one tool among many; limiting robots to only that would be a failure of imagination.

If we can build general intelligence that understands the underlying physics of interaction, then the shape of the hand becomes secondary to the intelligence that drives it. A suction pad, a gripper, a brush, a plasma welding nozzle – are all just different interfaces through which the same intelligence can reshape the physical world. The future of robot hands won’t look like ours. It will look more like a toolbox with a thousand hands: augmented, recombined, and scaled. Robots were always meant to extend what humans can do – to empower people to shape the physical world in places and at scales we never could before. Read more about our work at:

Great work! We have a low-cost, 3D-printed quick-swapping system build on a 1 DoF gripper that generalizes across everyday tools We can even attach it to a $7 hand :)

What about my excavator embodiment?

Gen-1 is Trillion time's better then figure toys

Amazing! We have our own quick change mechanism on our custom build bi-manual robot

Really cool stuff! One question, though. How does the model adjust the action space it outputs in real time when end effector changes? Does the system have to communicate to the model that an axis changed? For example, when gripper gets taped up.

finally an intern who can use both the browser and his hands

Smart move One brain, any hand. That's how you scale real robot dexterity.

More hands, more learning opportunities

One model, infinite hands 🤖 GEN-1 learning from all end effectors is huge

This is the future

If one policy can absorb a hand swap mid-rollout, embodiment starts looking like context rather than identity. The next leap is letting the robot choose and swap tools itself—turning hardware selection into inference-time reasoning.

single hand models are cool

The interesting metric isn’t how many end effectors it supports. It’s how much performance survives when you stop optimizing for just one.

OpenAI announced Project Camellia, a 3.2 GW data-center campus in Georgia. Construction is planned in phases from 2028 to 2032. OpenAI says it will cover the infrastructure costs, use closed-loop water systems, and fund $80M in local benefits. Those are firm commitments. Execution and local impact still need to be demonstrated. Source:

The many hands framing makes a lot of sense. Robots probably shouldn’t be forced into one human like hand design when different tasks need different tools.

This is awesome work

@huaijiangzhu So sick

that switch at ~1:00

@elonmusk buy them! 🙏🏼

incredible !!!

This is incredible. So I can apply this to robots I build? I am working on a cooking robot right now

@peteflorence akin to magic

Can startups use your models on there robots ??

Congratulations!

Very cool

GEN-1 learns from many, works with all

