Загрузка видео...

Не удалось загрузить видео

На главную

Our model can now learn from its own experience with RL! Our new π*0.6 model can more than double throughput over a base model trained without RL, and can perform real-world tasks: making espresso drinks, folding diverse laundry, and assembling boxes. More in the thread below.

710,973 просмотров • 10 месяцев назад •via X (Twitter)

Комментарии: 40

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

We train a general-purpose value function on all of our data, which tells the π*0.6 VLA which actions are good or bad. By asking π*0.6 to produce only good actions, we get better performance. We call this method Recap.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

π*0.6 can then collect more autonomous data, which can be used to further train the value function and further improve π*0.6! During autonomous data collection, a teleoperator can also intervene and provide corrections for significant mistakes, coaching π*0.6 further.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

We used π*0.6 to make various espresso drinks at the office. Here it runs for about 13 hours, with only a few interruptions.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

We also tested it out with diverse and realistic laundry items. Here π*0.6 folds laundry for 3 hours, averaging 3 minutes per item.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

We also trained π*0.6 to assemble boxes. Here is an hour of box building, with about two and a half minutes per box.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

Quantitatively, training π*0.6 with RL can more than double throughput (number of successful task executions per hour) on the hardest tasks and cut the number of failures by as much as a factor of two.

Фото профиля Physical Intelligence
Physical Intelligence10 месяцев назад

To learn more, see more videos, and read a full research paper about Recap and π*0.6, see our blog post here:

Фото профиля James Lin
James Lin10 месяцев назад

Love the 0% marketing, 100% what can the thing do

Фото профиля Shaun Maguire
Shaun Maguire10 месяцев назад

This is a sick way to dial in an espresso machine!

Фото профиля Deeptics x DeepHub
Deeptics x DeepHub10 месяцев назад

Really impressive direction combining vision-language-action with reinforcement learning is exactly what pushes robots closer to adaptive, real-world competence. On-the-job learning is a huge step for reducing hand-engineered behaviors and moving toward more generalizable, robust skill acquisition. Excited to see how these models evolve across broader task distributions.

Фото профиля clo
clo10 месяцев назад

0:32 did mr. robot just tap the edge of portafilter to get rid of air gaps 🥹

Фото профиля Olivia Li
Olivia Li10 месяцев назад

@donaldjewkes the GOAT makes another video

Фото профиля Cybernetic Labs
Cybernetic Labs10 месяцев назад

Strong results, especially the throughput gains from closing the loop with real-world RL. The next question is how the policy manages drift, recovery, and stability over long horizons across those task families.

Фото профиля Contextrix
Contextrix10 месяцев назад

Wow! Robots learning on their own. That’s cool!

Фото профиля A.I. NXTGEN Studio
A.I. NXTGEN Studio10 месяцев назад

This is so exciting!😁

Фото профиля Adam Patni
Adam Patni10 месяцев назад

Awesome work! Robots being described as decisive/indecisive is pretty funny - are there actual evals on this?

Фото профиля Wiz 👨‍🚀
Wiz 👨‍🚀10 месяцев назад

No latte art? 😋

Фото профиля berni
berni10 месяцев назад

let's go

Фото профиля nikhil ganesh
nikhil ganesh10 месяцев назад

Let’s gooo @donaldjewkes

Фото профиля daniel wu
daniel wu10 месяцев назад

/matcha latte😋

Фото профиля Sindre Reino Trosterud
Sindre Reino Trosterud10 месяцев назад

@ektrit these will make small businesses much more fun to run And yeah yeah you still need upstream a lot more

Фото профиля Luke Igel
Luke Igel10 месяцев назад

Let’s gooooo @donaldjewkes

Фото профиля OmnipotentCEO
OmnipotentCEO10 месяцев назад

As the Starbucks people continue to strike their way out of a job, they will be replaced by these incredible machines. I keep warning them but they say this tech does not exist smh people really have no idea what is coming.

Фото профиля avi
avi10 месяцев назад

congrats!

Фото профиля Nikhil Kulkarni
Nikhil Kulkarni10 месяцев назад

Loved the music choice in this piece. It helped me slow down and I appreciated the longer held shots so much more! I find it so interesting that the robot makes very human like choices like moving the glass 90% of where it needs to be and then pushing it a little more for that final push! Fantastic work!!

Фото профиля Eddy Xu
Eddy Xu10 месяцев назад

very very exciting

Фото профиля bodhi shartner
bodhi shartner10 месяцев назад

Whoop

Фото профиля Ed Henderson
Ed Henderson10 месяцев назад

It feels like scaling RL will depend on scene / environment resets at scale. It'd be funny if the unlock is teaching the robot to undo it's work? Or automating arm farm for scene resets?!

Фото профиля PAULORIZED
PAULORIZED10 месяцев назад

faster than the kids at my local Starbucks.

Фото профиля Lazarz
Lazarz10 месяцев назад

@dvruette looks like something we need

Фото профиля this is my least favorite life
this is my least favorite life10 месяцев назад

send this to starbucks workers on strike

Фото профиля 쇼팽.
쇼팽.10 месяцев назад

pi 0.61 next? 😂 A great build up for pi 3.14.

Фото профиля WeeklyRobotics
WeeklyRobotics10 месяцев назад

WHAT IS THIS SORCERY

Фото профиля Andres
Andres10 месяцев назад

This is sick

Фото профиля Vipul Divyanshu⚡
Vipul Divyanshu⚡10 месяцев назад

Exciting progress 🚀

Фото профиля Vihaan Shah
Vihaan Shah9 месяцев назад

This is cool, but the real puzzle is why has PI hired some great hardware engineers?

Фото профиля Cris Lenta
Cris Lenta10 месяцев назад

this is so impressive! congrats to the team!!

Фото профиля 16VC
16VC10 месяцев назад

The combination of autonomous data collection with human correction is a smart balance. Doubling throughput on real-world tasks shows how fast embodied AI is maturing.

Фото профиля Yash More
Yash More10 месяцев назад

this is beautiful. are the trajectories for tele-op opensourced? can i see the coffee and folding ones?

Фото профиля aadithva
aadithva9 месяцев назад

Unnecessarily cool. Future office approved. What’s the price tag on this beauty?

Похожие видео