正在加载视频...

视频加载失败

Our model can now learn from its own experience with RL! Our new π*0.6 model can more than double throughput over a base model trained without RL, and can perform real-world tasks: making espresso drinks, folding diverse laundry, and assembling boxes. More in the thread below.

710,973 次观看 • 10 个月前 •via X (Twitter)

40 条评论

Physical Intelligence 的头像
Physical Intelligence10 个月前

We train a general-purpose value function on all of our data, which tells the π*0.6 VLA which actions are good or bad. By asking π*0.6 to produce only good actions, we get better performance. We call this method Recap.

Physical Intelligence 的头像
Physical Intelligence10 个月前

π*0.6 can then collect more autonomous data, which can be used to further train the value function and further improve π*0.6! During autonomous data collection, a teleoperator can also intervene and provide corrections for significant mistakes, coaching π*0.6 further.

Physical Intelligence 的头像
Physical Intelligence10 个月前

We used π*0.6 to make various espresso drinks at the office. Here it runs for about 13 hours, with only a few interruptions.

Physical Intelligence 的头像
Physical Intelligence10 个月前

We also tested it out with diverse and realistic laundry items. Here π*0.6 folds laundry for 3 hours, averaging 3 minutes per item.

Physical Intelligence 的头像
Physical Intelligence10 个月前

We also trained π*0.6 to assemble boxes. Here is an hour of box building, with about two and a half minutes per box.

Physical Intelligence 的头像
Physical Intelligence10 个月前

Quantitatively, training π*0.6 with RL can more than double throughput (number of successful task executions per hour) on the hardest tasks and cut the number of failures by as much as a factor of two.

Physical Intelligence 的头像
Physical Intelligence10 个月前

To learn more, see more videos, and read a full research paper about Recap and π*0.6, see our blog post here:

James Lin 的头像
James Lin10 个月前

Love the 0% marketing, 100% what can the thing do

Shaun Maguire 的头像
Shaun Maguire10 个月前

This is a sick way to dial in an espresso machine!

Deeptics x DeepHub 的头像
Deeptics x DeepHub10 个月前

Really impressive direction combining vision-language-action with reinforcement learning is exactly what pushes robots closer to adaptive, real-world competence. On-the-job learning is a huge step for reducing hand-engineered behaviors and moving toward more generalizable, robust skill acquisition. Excited to see how these models evolve across broader task distributions.

clo 的头像
clo10 个月前

0:32 did mr. robot just tap the edge of portafilter to get rid of air gaps 🥹

Olivia Li 的头像
Olivia Li10 个月前

@donaldjewkes the GOAT makes another video

Cybernetic Labs 的头像
Cybernetic Labs10 个月前

Strong results, especially the throughput gains from closing the loop with real-world RL. The next question is how the policy manages drift, recovery, and stability over long horizons across those task families.

Contextrix 的头像
Contextrix10 个月前

Wow! Robots learning on their own. That’s cool!

A.I. NXTGEN Studio 的头像
A.I. NXTGEN Studio10 个月前

This is so exciting!😁

Adam Patni 的头像
Adam Patni10 个月前

Awesome work! Robots being described as decisive/indecisive is pretty funny - are there actual evals on this?

Wiz 👨‍🚀 的头像
Wiz 👨‍🚀10 个月前

No latte art? 😋

berni 的头像
berni10 个月前

let's go

nikhil ganesh 的头像
nikhil ganesh10 个月前

Let’s gooo @donaldjewkes

daniel wu 的头像
daniel wu10 个月前

/matcha latte😋

Sindre Reino Trosterud 的头像
Sindre Reino Trosterud10 个月前

@ektrit these will make small businesses much more fun to run And yeah yeah you still need upstream a lot more

Luke Igel 的头像
Luke Igel10 个月前

Let’s gooooo @donaldjewkes

OmnipotentCEO 的头像
OmnipotentCEO10 个月前

As the Starbucks people continue to strike their way out of a job, they will be replaced by these incredible machines. I keep warning them but they say this tech does not exist smh people really have no idea what is coming.

avi 的头像
avi10 个月前

congrats!

Nikhil Kulkarni 的头像
Nikhil Kulkarni10 个月前

Loved the music choice in this piece. It helped me slow down and I appreciated the longer held shots so much more! I find it so interesting that the robot makes very human like choices like moving the glass 90% of where it needs to be and then pushing it a little more for that final push! Fantastic work!!

Eddy Xu 的头像
Eddy Xu10 个月前

very very exciting

bodhi shartner 的头像
bodhi shartner10 个月前

Whoop

Ed Henderson 的头像
Ed Henderson10 个月前

It feels like scaling RL will depend on scene / environment resets at scale. It'd be funny if the unlock is teaching the robot to undo it's work? Or automating arm farm for scene resets?!

PAULORIZED 的头像
PAULORIZED10 个月前

faster than the kids at my local Starbucks.

Lazarz 的头像
Lazarz10 个月前

@dvruette looks like something we need

this is my least favorite life 的头像
this is my least favorite life10 个月前

send this to starbucks workers on strike

쇼팽. 的头像
쇼팽.10 个月前

pi 0.61 next? 😂 A great build up for pi 3.14.

WeeklyRobotics 的头像
WeeklyRobotics10 个月前

WHAT IS THIS SORCERY

Andres 的头像
Andres10 个月前

This is sick

Vipul Divyanshu⚡ 的头像
Vipul Divyanshu⚡10 个月前

Exciting progress 🚀

Vihaan Shah 的头像
Vihaan Shah9 个月前

This is cool, but the real puzzle is why has PI hired some great hardware engineers?

Cris Lenta 的头像
Cris Lenta10 个月前

this is so impressive! congrats to the team!!

16VC 的头像
16VC10 个月前

The combination of autonomous data collection with human correction is a smart balance. Doubling throughput on real-world tasks shows how fast embodied AI is maturing.

Yash More 的头像
Yash More10 个月前

this is beautiful. are the trajectories for tele-op opensourced? can i see the coffee and folding ones?

aadithva 的头像
aadithva9 个月前

Unnecessarily cool. Future office approved. What’s the price tag on this beauty?

相关视频