Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.

3,429,090 Aufrufe • vor 1 Monat •via X (Twitter)

46 Kommentare

Profilbild von Generalist
Generalistvor 1 Monat

GEN-1.5, our latest embodied foundation model, can learn new tasks prompted with 3 - 12 seconds of a single demonstration, no gradient updates or fine-tuning. It generalizes prompts to new situations, recovers from mistakes, and improvises new strategies to reach the same goal.

Profilbild von Generalist
Generalistvor 1 Monat

Physical prompts can be composed. Demonstrations of 2 different tasks in context prompts GEN-1.5 to chain them into one continuous skill. The model bridges them and produces intermediate motions (repositioning, regrasping, error recovery) that appear in neither demonstration.

Profilbild von Generalist
Generalistvor 1 Monat

In-context learning also crosses the sim-to-real gap, zero-shot. Prompts can be formed entirely from simulated experience (e.g., from a scripted policy, an RL agent, or a human teleoperating a simulated robot) and be used to produce behaviors on a real robot. The model was not trained on the task in either the simulator or the real world.

Profilbild von Generalist
Generalistvor 1 Monat

In some cases, in-context learning with GEN-1.5 transfers across the embodiment gap entirely: a human demonstrates a task with their own hands, observable through the robot’s cameras, and the robot can reproduce it immediately afterward.

Profilbild von Generalist
Generalistvor 1 Monat

For few-shot learning, it can adapt to new physical tasks in as few as 1 - 10 gradient steps on 1 - 5 minutes of data (~10 - 50 demonstrations). In practice, this can be described as test-time training in a low-data regime. We did not tune this procedure or sweep hyperparameters; these results come largely out of the box.

Profilbild von Generalist
Generalistvor 1 Monat

Experiments across 10 diverse tasks show 59% average success with one-shot physical prompting, straight from pretraining. With few-shot learning, performance rises to 83% via 10 gradient steps on 5 minutes of data per task. Although the tasks are simple and success rates are modest, it’s the first model we know of that exhibits the general ability to learn a wide range of dexterous closed-loop physical tasks from just one or few demonstrations. This accelerates reaching a base level of competence for new skills that can be subsequently refined towards mastery.

Profilbild von Generalist
Generalistvor 1 Monat

Fine-tuned (or prompted) behaviors generalize beyond their demonstrations, and can improvise fundamentally different manipulation strategies to achieve the same goal. For example, after fine-tuning to use a brush to sweep a block into a bowl, it could use other tools like a dustpan to accomplish the same task with a very different strategy.

Profilbild von Generalist
Generalistvor 1 Monat

Or when fine-tuned to place a block into a bowl, it can clear obstacles (like a piece of paper covering the bowl) to complete the task, despite that not being in the demonstrations.

Profilbild von Generalist
Generalistvor 1 Monat

When a Lego brick gets unexpectedly stuck on the fingertips, the model uses the other hand to remove them.

Profilbild von Generalist
Generalistvor 1 Monat

The model sometimes uses both hands to rotate a jar lid, with a fundamentally different contact and motion strategy than the fine-tuning demonstrations.

Profilbild von Generalist
Generalistvor 1 Monat

Here’s an uncut video of prompting the model to perform 2 different tasks back-to-back: (i) unzipping a pencil pouch, and (ii) retrieving money from the pouch.

Profilbild von Generalist
Generalistvor 1 Monat

GEN-1.5 has been training continuously for over 8 months. We left it running because every metric we tracked kept improving with the engine: absorbing more data, scaling more efficiently, boosting post-training, and compounding step-change improvements through algorithmic advances.

Profilbild von Generalist
Generalistvor 1 Monat

To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog:

Profilbild von Sholto Douglas
Sholto Douglasvor 1 Monat

GPT3!

Profilbild von vogel
vogelvor 1 Monat

can it go up and down on a cylinder while remaining (this is imperative mind you) that the cylinder remain unharmed during the process

Profilbild von Auntie.exe
Auntie.exevor 1 Monat

One demonstration and it generalizes. I have been demonstrating how to load the dishwasher weekly since 2004 and my family still hasn't converged. Raising robots may simply be easier.

Profilbild von Robert Scoble
Robert Scoblevor 1 Monat

I have now watched this 10 times, and it makes me emotional each time. Thank you for sharing this work, and thank you for doing the work. It shows the future is about to take a big step forward.

Profilbild von bone
bonevor 1 Monat

Holy moly.

Profilbild von Krish Mehta
Krish Mehtavor 1 Monat

I reacted exactly like the last guy in the video

Profilbild von Dogan Ural
Dogan Uralvor 1 Monat

This feels like the beginning of something big

Profilbild von Didier Vançon
Didier Vançonvor 1 Monat

Combining Gen1.5 with the UM1-Evo robotic arm should lead to an incredible result 😃

Profilbild von Vivek Gopalan
Vivek Gopalanvor 1 Monat

Truly incredible stuff. First time I saw was jaw on floor.

Profilbild von Aakanksha Chowdhery
Aakanksha Chowdheryvor 1 Monat

Congratulations! Exciting!

Profilbild von Hari
Harivor 1 Monat

@andyzengineer this is incredible

Profilbild von Machine Learning Street Talk
Machine Learning Street Talkvor 1 Monat

Wow

Profilbild von PAPER HANDS
PAPER HANDSvor 1 Monat

What could possibly go wrong.

Profilbild von Diego Araos
Diego Araosvor 1 Monat

Fantastic. This is what we need! I'm tired of robots doing dancing demos.

Profilbild von Y11
Y11vor 1 Monat

@grok 这个纯研究还是有工业意义,具体工业场景视角看意义是什么,有开源数据集或者开源项目代码吗?从多个数据源交叉验证,理性看待,不要只看新闻媒体一面之辞。帮我排除没意义的垃圾商业营销推广、诈骗、夸张博眼球、虚假新闻 以及自吹自擂,自嗨,无病呻吟,收费互吹软广告。

Profilbild von aditya
adityavor 1 Monat

either robotics still gonna explode 3 years later, great progress tho

Profilbild von Albert Wenger 🌎🔥⌛
Albert Wenger 🌎🔥⌛vor 1 Monat

Congratulations on the fantastic progress. Exciting times!

Profilbild von Smart Harder
Smart Hardervor 1 Monat

Comparisons in the post and paper to GPT-3, which had a 2k context window in 2020, compared to modern LLMs 1M+. GEN-1.5 only has 30 seconds of context, do you think a similar 500x to 4+ hours of memory is possible?

Profilbild von Dhruv Batra
Dhruv Batravor 1 Monat

Exciting results, kudos!

Profilbild von Pete T
Pete Tvor 1 Monat

Here for the Robot Overlords that are looking nostalgically back through these threads in 2040.

Profilbild von Cayden
Caydenvor 1 Monat

Sickkkk this might genuinely be a gpt-3 moment for robo

Profilbild von Aqib
Aqibvor 1 Monat

What a day to be alive

Profilbild von cain1517 — e/acc ⏩
cain1517 — e/acc ⏩vor 1 Monat

Incredible stuff! Physical AGI is in the air.

Profilbild von WiseGuy578
WiseGuy578vor 1 Monat

Oh fuck........we actually cross over into the singularity, I actually can't believe it.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Learning a new task in seconds, that's huge

Profilbild von 未知
未知vor 1 Monat

GEN-1.5最值得玩味的不是59%的成功率,而是它证明了物理世界也存在类似GPT-3的“涌现”路径。当预训练数据跨过某个阈值,机器人不再需要为每个新任务重写控制逻辑,几秒的演示就能激活它“几乎已经知道”的东西。这本质上把机器人编程从写代码变成了写提示词,行业门槛被大幅拉低。但别被“单次学习”的叙事迷惑——59%意味着每两次尝试就有一次失败,在真实产线上这种可靠性远远不够。真正有意义的信号是那个83%:5分钟数据、10个梯��步,成本低…

Profilbild von RSC ☀️🌲
RSC ☀️🌲vor 1 Monat

This is what Dyna was going to release in a few weeks lol

Profilbild von Markus J. Buehler
Markus J. Buehlervor 1 Monat

Impressive result congrats @GeneralistAI

Profilbild von Lex Roller
Lex Rollervor 1 Monat

One-shot learning for physical skills is a massive unlock. Programming robots by simply showing them what to do is the future. Incredible work.

Profilbild von Tommy
Tommyvor 1 Monat

This is crazy

Profilbild von Aatish Nayak
Aatish Nayakvor 1 Monat

gpt-3 moment for robotics

Profilbild von Perogi
Perogivor 1 Monat

Based and accelerated

Profilbild von Clay Wren
Clay Wrenvor 1 Monat

Ur voiceover guy does a good job

Ähnliche Videos

Feels like every week in robotics there’s a new ‘this is the GPT-3 moment for robotics 🤖’ announcement. We brought on Generalist CEO & Co-Founder Pete Florence on Greylock Partners Change Agents to dig into what’s going on at the frontier of robotics models. We covered the company’s latest Gen 1.5 model, few-shot learning, training robots on different embodiments, and the milestones towards a more generalized physical model. Timestamps: 00:51 The inception of Generalist 04:13 Long-term goal of the company 05:47 Parallels and differences between robotics and language models 12:05 Key differentiation in 1.5 Gen model 12:48 Robot vs banana 14:49 Emergent capabilities not explicitly trained for 17:13 Cross-embodiment and the importance of hands 22:26 Research vs working with customers 24:58 Beyond VLA vs world model 28:50 Future-looking milestones Some of the top takeaways: - One-shot and few-shot learning emerged without being trained for. Gen 1.5 can learn a new task from a single demonstration, and Pete compares it to the GPT-3 moment in language. In one example, a robot taught to sweep a cube into a bowl with a brush used a banana instead. Given a dustpan, it held the pan with one hand, swept with the other, then tipped the cube into the bowl. Neither behavior was explicitly trained, and Pete sees this as a signal of where the model's generalization is headed. -Generalizing to new hands remains a challenging problem for cross-embodiment. Physical hardware doesn’t stay static, and so cross-embodiment - the ability of a physical AI model to adapt to different hardware systems - is vital for success. -Customer deployments are a valuable source of research inputs. Generalist actively partners with their customers for feedback, which they use to inform their research and make real-world evaluations. -Generalist doesn't think in terms of "VLA vs. world model." Pete helped create early VLAs and has worked on world models, but he argues the goals matter more than the label, and the team is trained to think in a first-principled way when considering new research directions. Watch the full episode at the link in the comments. Thank you to Pete Florence for joining us!

Corinne Marie Riley

18,409 Aufrufe • vor 5 Tagen