Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Our model can now learn from its own experience with RL! Our new π*0.6 model can more than double throughput over a base model trained without RL, and can perform real-world tasks: making espresso drinks, folding diverse laundry, and assembling boxes. More in the thread below.

710,973 Aufrufe • vor 10 Monaten •via X (Twitter)

40 Kommentare

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

We train a general-purpose value function on all of our data, which tells the π*0.6 VLA which actions are good or bad. By asking π*0.6 to produce only good actions, we get better performance. We call this method Recap.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

π*0.6 can then collect more autonomous data, which can be used to further train the value function and further improve π*0.6! During autonomous data collection, a teleoperator can also intervene and provide corrections for significant mistakes, coaching π*0.6 further.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

We used π*0.6 to make various espresso drinks at the office. Here it runs for about 13 hours, with only a few interruptions.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

We also tested it out with diverse and realistic laundry items. Here π*0.6 folds laundry for 3 hours, averaging 3 minutes per item.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

We also trained π*0.6 to assemble boxes. Here is an hour of box building, with about two and a half minutes per box.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

Quantitatively, training π*0.6 with RL can more than double throughput (number of successful task executions per hour) on the hardest tasks and cut the number of failures by as much as a factor of two.

Profilbild von Physical Intelligence
Physical Intelligencevor 10 Monaten

To learn more, see more videos, and read a full research paper about Recap and π*0.6, see our blog post here:

Profilbild von James Lin
James Linvor 10 Monaten

Love the 0% marketing, 100% what can the thing do

Profilbild von Shaun Maguire
Shaun Maguirevor 10 Monaten

This is a sick way to dial in an espresso machine!

Profilbild von Deeptics x DeepHub
Deeptics x DeepHubvor 10 Monaten

Really impressive direction combining vision-language-action with reinforcement learning is exactly what pushes robots closer to adaptive, real-world competence. On-the-job learning is a huge step for reducing hand-engineered behaviors and moving toward more generalizable, robust skill acquisition. Excited to see how these models evolve across broader task distributions.

Profilbild von clo
clovor 10 Monaten

0:32 did mr. robot just tap the edge of portafilter to get rid of air gaps 🥹

Profilbild von Olivia Li
Olivia Livor 10 Monaten

@donaldjewkes the GOAT makes another video

Profilbild von Cybernetic Labs
Cybernetic Labsvor 10 Monaten

Strong results, especially the throughput gains from closing the loop with real-world RL. The next question is how the policy manages drift, recovery, and stability over long horizons across those task families.

Profilbild von Contextrix
Contextrixvor 10 Monaten

Wow! Robots learning on their own. That’s cool!

Profilbild von A.I. NXTGEN Studio
A.I. NXTGEN Studiovor 10 Monaten

This is so exciting!😁

Profilbild von Adam Patni
Adam Patnivor 10 Monaten

Awesome work! Robots being described as decisive/indecisive is pretty funny - are there actual evals on this?

Profilbild von Wiz 👨‍🚀
Wiz 👨‍🚀vor 10 Monaten

No latte art? 😋

Profilbild von berni
bernivor 10 Monaten

let's go

Profilbild von nikhil ganesh
nikhil ganeshvor 10 Monaten

Let’s gooo @donaldjewkes

Profilbild von daniel wu
daniel wuvor 10 Monaten

/matcha latte😋

Profilbild von Sindre Reino Trosterud
Sindre Reino Trosterudvor 10 Monaten

@ektrit these will make small businesses much more fun to run And yeah yeah you still need upstream a lot more

Profilbild von Luke Igel
Luke Igelvor 10 Monaten

Let’s gooooo @donaldjewkes

Profilbild von OmnipotentCEO
OmnipotentCEOvor 10 Monaten

As the Starbucks people continue to strike their way out of a job, they will be replaced by these incredible machines. I keep warning them but they say this tech does not exist smh people really have no idea what is coming.

Profilbild von avi
avivor 10 Monaten

congrats!

Profilbild von Nikhil Kulkarni
Nikhil Kulkarnivor 10 Monaten

Loved the music choice in this piece. It helped me slow down and I appreciated the longer held shots so much more! I find it so interesting that the robot makes very human like choices like moving the glass 90% of where it needs to be and then pushing it a little more for that final push! Fantastic work!!

Profilbild von Eddy Xu
Eddy Xuvor 10 Monaten

very very exciting

Profilbild von bodhi shartner
bodhi shartnervor 10 Monaten

Whoop

Profilbild von Ed Henderson
Ed Hendersonvor 10 Monaten

It feels like scaling RL will depend on scene / environment resets at scale. It'd be funny if the unlock is teaching the robot to undo it's work? Or automating arm farm for scene resets?!

Profilbild von PAULORIZED
PAULORIZEDvor 10 Monaten

faster than the kids at my local Starbucks.

Profilbild von Lazarz
Lazarzvor 10 Monaten

@dvruette looks like something we need

Profilbild von this is my least favorite life
this is my least favorite lifevor 10 Monaten

send this to starbucks workers on strike

Profilbild von 쇼팽.
쇼팽.vor 10 Monaten

pi 0.61 next? 😂 A great build up for pi 3.14.

Profilbild von WeeklyRobotics
WeeklyRoboticsvor 10 Monaten

WHAT IS THIS SORCERY

Profilbild von Andres
Andresvor 10 Monaten

This is sick

Profilbild von Vipul Divyanshu⚡
Vipul Divyanshu⚡vor 10 Monaten

Exciting progress 🚀

Profilbild von Vihaan Shah
Vihaan Shahvor 9 Monaten

This is cool, but the real puzzle is why has PI hired some great hardware engineers?

Profilbild von Cris Lenta
Cris Lentavor 10 Monaten

this is so impressive! congrats to the team!!

Profilbild von 16VC
16VCvor 10 Monaten

The combination of autonomous data collection with human correction is a smart balance. Doubling throughput on real-world tasks shows how fast embodied AI is maturing.

Profilbild von Yash More
Yash Morevor 10 Monaten

this is beautiful. are the trajectories for tele-op opensourced? can i see the coffee and folding ones?

Profilbild von aadithva
aadithvavor 9 Monaten

Unnecessarily cool. Future office approved. What’s the price tag on this beauty?

Ähnliche Videos