Загрузка видео...
Не удалось загрузить видео
Our model can now learn from its own experience with RL! Our new π*0.6 model can more than double throughput over a base model trained without RL, and can perform real-world tasks: making espresso drinks, folding diverse laundry, and assembling boxes. More in the thread below.
710,973 просмотров • 10 месяцев назад •via X (Twitter)
Комментарии: 40

We train a general-purpose value function on all of our data, which tells the π*0.6 VLA which actions are good or bad. By asking π*0.6 to produce only good actions, we get better performance. We call this method Recap.

π*0.6 can then collect more autonomous data, which can be used to further train the value function and further improve π*0.6! During autonomous data collection, a teleoperator can also intervene and provide corrections for significant mistakes, coaching π*0.6 further.

We used π*0.6 to make various espresso drinks at the office. Here it runs for about 13 hours, with only a few interruptions.

We also tested it out with diverse and realistic laundry items. Here π*0.6 folds laundry for 3 hours, averaging 3 minutes per item.

We also trained π*0.6 to assemble boxes. Here is an hour of box building, with about two and a half minutes per box.

Quantitatively, training π*0.6 with RL can more than double throughput (number of successful task executions per hour) on the hardest tasks and cut the number of failures by as much as a factor of two.

To learn more, see more videos, and read a full research paper about Recap and π*0.6, see our blog post here:

Love the 0% marketing, 100% what can the thing do

This is a sick way to dial in an espresso machine!

Really impressive direction combining vision-language-action with reinforcement learning is exactly what pushes robots closer to adaptive, real-world competence. On-the-job learning is a huge step for reducing hand-engineered behaviors and moving toward more generalizable, robust skill acquisition. Excited to see how these models evolve across broader task distributions.

0:32 did mr. robot just tap the edge of portafilter to get rid of air gaps 🥹

@donaldjewkes the GOAT makes another video

Strong results, especially the throughput gains from closing the loop with real-world RL. The next question is how the policy manages drift, recovery, and stability over long horizons across those task families.

Wow! Robots learning on their own. That’s cool!

This is so exciting!😁

Awesome work! Robots being described as decisive/indecisive is pretty funny - are there actual evals on this?

No latte art? 😋

let's go

Let’s gooo @donaldjewkes

/matcha latte😋

@ektrit these will make small businesses much more fun to run And yeah yeah you still need upstream a lot more

Let’s gooooo @donaldjewkes

As the Starbucks people continue to strike their way out of a job, they will be replaced by these incredible machines. I keep warning them but they say this tech does not exist smh people really have no idea what is coming.

congrats!

Loved the music choice in this piece. It helped me slow down and I appreciated the longer held shots so much more! I find it so interesting that the robot makes very human like choices like moving the glass 90% of where it needs to be and then pushing it a little more for that final push! Fantastic work!!

very very exciting

Whoop

It feels like scaling RL will depend on scene / environment resets at scale. It'd be funny if the unlock is teaching the robot to undo it's work? Or automating arm farm for scene resets?!

faster than the kids at my local Starbucks.

@dvruette looks like something we need

send this to starbucks workers on strike

pi 0.61 next? 😂 A great build up for pi 3.14.

WHAT IS THIS SORCERY

This is sick

Exciting progress 🚀

This is cool, but the real puzzle is why has PI hired some great hardware engineers?

this is so impressive! congrats to the team!!

The combination of autonomous data collection with human correction is a smart balance. Doubling throughput on real-world tasks shows how fast embodied AI is maturing.

this is beautiful. are the trajectories for tele-op opensourced? can i see the coffee and folding ones?

Unnecessarily cool. Future office approved. What’s the price tag on this beauty?
