Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.

301,754 görüntüleme • 1 ay önce •via X (Twitter)

32 Yorum

Generalist profil fotoğrafı
Generalist1 ay önce

Each hand is a different sensorimotor interface by which GEN-1 experiences the physical world. Scaling pretraining across thousands of these interfaces teaches GEN-1 a universal physical commonsense that transfers to new hands and new ways to grasp, push, pull, twist, and more.

Generalist profil fotoğrafı
Generalist1 ay önce

Every hand is a different language for physical interaction. Just as training on multiple languages produces more capable LLMs, a model trained across embodiments can gather shared knowledge across instances and more readily separate what is specific to a hand from what is universal about the world.

Generalist profil fotoğrafı
Generalist1 ay önce

Unlike humans, robots aren’t locked into the hands they’re born with. Machines can easily swap mechanical hands, and tool changers are common in automation. Switching between different end effectors to reach a goal can be viewed as a form of physical reasoning: using the right tool for the right job, much as multilingual chain-of-thought can improve LLM reasoning for downstream RL.

Generalist profil fotoğrafı
Generalist1 ay önce

What happens when you change the hand mid-task? We tested this by modifying the hands mid-rollout and letting the same model keep running. It perceives the new tool, and finds a new trajectory and contact strategy to complete the task. This works because training on mixed data forces the model to condition its behavior on the hand in front of it. This brings it closer to a general understanding of how shapes and contact surfaces interact with the physical world – and which actuation strategy should follow.

Generalist profil fotoğrafı
Generalist1 ay önce

Nature didn’t converge on just a single solution for manipulating the physical world. It exploded into millions. You can see this everywhere in the diversity of life around us: from the beak of a bird, to the trunk of an elephant, from the suction pads of an octopus, to the pollen baskets of a honeybee. The human hand is only a single point in a design space so large we’ve barely begun to map. For machines, five-fingered hands will just be one tool among many; limiting robots to only that would be a failure of imagination.

Generalist profil fotoğrafı
Generalist1 ay önce

If we can build general intelligence that understands the underlying physics of interaction, then the shape of the hand becomes secondary to the intelligence that drives it. A suction pad, a gripper, a brush, a plasma welding nozzle – are all just different interfaces through which the same intelligence can reshape the physical world. The future of robot hands won’t look like ours. It will look more like a toolbox with a thousand hands: augmented, recombined, and scaled. Robots were always meant to extend what humans can do – to empower people to shape the physical world in places and at scales we never could before. Read more about our work at:

Xianyi Cheng profil fotoğrafı
Xianyi Cheng1 ay önce

Great work! We have a low-cost, 3D-printed quick-swapping system build on a 1 DoF gripper that generalizes across everyday tools We can even attach it to a $7 hand :)

Lane Burgett profil fotoğrafı
Lane Burgett1 ay önce

What about my excavator embodiment?

maru profil fotoğrafı
maru1 ay önce

Gen-1 is Trillion time's better then figure toys

Mikhail profil fotoğrafı
Mikhail1 ay önce

Amazing! We have our own quick change mechanism on our custom build bi-manual robot

John Dagdelen profil fotoğrafı
John Dagdelen1 ay önce

Really cool stuff! One question, though. How does the model adjust the action space it outputs in real time when end effector changes? Does the system have to communicate to the model that an axis changed? For example, when gripper gets taped up.

Clawbrowser - Anti-Detect Browser for AI Agents profil fotoğrafı
Clawbrowser - Anti-Detect Browser for AI Agents1 ay önce

finally an intern who can use both the browser and his hands

Ella Tech & Tool profil fotoğrafı
Ella Tech & Tool1 ay önce

Smart move One brain, any hand. That's how you scale real robot dexterity.

Zane profil fotoğrafı
Zane1 ay önce

More hands, more learning opportunities

Bblack profil fotoğrafı
Bblack1 ay önce

One model, infinite hands 🤖 GEN-1 learning from all end effectors is huge

Akriti profil fotoğrafı
Akriti1 ay önce

This is the future

Emma profil fotoğrafı
Emma1 ay önce

If one policy can absorb a hand swap mid-rollout, embodiment starts looking like context rather than identity. The next leap is letting the robot choose and swap tools itself—turning hardware selection into inference-time reasoning.

SEAR profil fotoğrafı
SEAR1 ay önce

single hand models are cool

Context profil fotoğrafı
Context1 ay önce

The interesting metric isn’t how many end effectors it supports. It’s how much performance survives when you stop optimizing for just one.

Mi.lu. profil fotoğrafı
Mi.lu.1 ay önce

OpenAI announced Project Camellia, a 3.2 GW data-center campus in Georgia. Construction is planned in phases from 2028 to 2032. OpenAI says it will cover the infrastructure costs, use closed-loop water systems, and fund $80M in local benefits. Those are firm commitments. Execution and local impact still need to be demonstrated. Source:

Lex profil fotoğrafı
Lex1 ay önce

The many hands framing makes a lot of sense. Robots probably shouldn’t be forced into one human like hand design when different tasks need different tools.

Guillermo Duró profil fotoğrafı
Guillermo Duró1 ay önce

This is awesome work

Nihal profil fotoğrafı
Nihal1 ay önce

@huaijiangzhu So sick

Andy Wojcicki profil fotoğrafı
Andy Wojcicki1 ay önce

that switch at ~1:00

ArtuR2 ⏩️⤴️ profil fotoğrafı
ArtuR2 ⏩️⤴️1 ay önce

@elonmusk buy them! 🙏🏼

Bela Becerra profil fotoğrafı
Bela Becerra1 ay önce

incredible !!!

Sam W profil fotoğrafı
Sam W1 ay önce

This is incredible. So I can apply this to robots I build? I am working on a cooking robot right now

Vivek Gopalan profil fotoğrafı
Vivek Gopalan1 ay önce

@peteflorence akin to magic

Ramzi RAMZI profil fotoğrafı
Ramzi RAMZI1 ay önce

Can startups use your models on there robots ??

Han profil fotoğrafı
Han1 ay önce

Congratulations!

Julian Fried profil fotoğrafı
Julian Fried1 ay önce

Very cool

Noor Tech profil fotoğrafı
Noor Tech1 ay önce

GEN-1 learns from many, works with all

Benzer Videolar

We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We discovered a near-perfect log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and this loss directly predicts real-robot success rate. Humanoid robots will be the end game, because they are the practical form factor with minimal embodiment gap from humans. Call it the Bitter Lesson of robot hardware: the kinematic similarity lets us simply retarget human finger motion onto dexterous robot hand joints. No learned embeddings, no fancy transfer algorithms needed. Relative wrist motion + retargeted 22-DoF finger actions serve as a unified action space that carries through from pre-training to robot execution. Our recipe is called "EgoScale": - Pre-train GR00T N1.5 on 20K hours of human video, mid-train with only 4 hours (!) of robot play data with Sharpa hands. 54% gains over training from scratch across 5 highly dexterous tasks. - Most surprising result: a *single* teleop demo is sufficient to learn a never-before-seen task. Our recipe enables extreme data efficiency. - Although we pre-train in 22-DoF hand joint space, the policy transfers to a Unitree G1 with 7-DoF tri-finger hands. 30%+ gains over training on G1 data alone. The scalable path to robot dexterity was never more robots. It was always us. Deep dives in thread:

Jim Fan

300,427 görüntüleme • 6 ay önce