正在加载视频...

视频加载失败

Today we’re launching Enact: post-training infrastructure that makes robotics models work in the real world. Robotics models fail when execution reaches states their training data never covered. A slip, an off-angle grasp, or an external disturbance can leave the policy without a learned recovery. Enact finds those failures and...

52,909 次观看 • 1 个月前 •via X (Twitter)

62 条评论

Kyle Vedder 的头像
Kyle Vedder1 个月前

Exciting release! We need more data groups that deeply understand DAgger

James Stevens 的头像
James Stevens1 个月前

Thanks Kyle :)

James Stevens 的头像
James Stevens1 个月前

To isolate why recovery data matters, we used a controlled packing task common in e-commerce fulfillment, with a single item. We collected 200 demonstrations of successful placements into the box (the “happy path”) and fine-tuned π0.5. The policy failed on 10 of 100 rollouts. Adding more happy-path demonstrations did not improve success rates. Targeted recovery data did.

James Stevens 的头像
James Stevens1 个月前

Why? Happy-path demonstrations do not cover the out-of-distribution states created by the policy’s own mistakes. Every example showed a successful handoff, so the policy received no training signal for recovering after a drop.

James Stevens 的头像
James Stevens1 个月前

After adding 50 targeted recovery demonstrations, the retrained policy succeeded on 99 of 100 consecutive packing rollouts. It recovered from every table drop.

James Stevens 的头像
James Stevens1 个月前

This is a simple application of Dataset Aggregation (DAgger): roll out the current policy, collect targeted recovery demonstrations at the failure states it actually reaches, aggregate them into the training set, and retrain. As task complexity increases, failure modes multiply, rare failures become harder to discover, and collecting the recovery data needed to solve them takes substantially more effort.

James Stevens 的头像
James Stevens1 个月前

In our table-bussing task, the obvious failures were cups or plates flipping. Those were easy to predict and collect recovery data for. But roughly once every 100 runs, the knife wedged itself under the bin.

James Stevens 的头像
James Stevens1 个月前

The only way to recover was to move the bin itself, which meant training the robot on an entirely new skill! This is the long tail that breaks deployed systems: a rare failure can demand a behavior the original task never required. Finding these states and generating the data to solve them is what Enact does.

James Stevens 的头像
James Stevens1 个月前

We’re already serving our first customers and applying this loop to failures their models encounter in the real world. Deploying a policy that fails on a real task? Tell us the model, task, and where it breaks: [email protected] or DM me.

Lyron 的头像
Lyron1 个月前

This is awesome

Ethan Breitkreutz 的头像
Ethan Breitkreutz1 个月前

@ycombinator This is really interesting. Excited to see the progression as you guys continue forward.

James Stevens 的头像
James Stevens1 个月前

@ethan_breitk @ycombinator Thanks Ethan! Stay tuned ;)

Ethan Breitkreutz 的头像
Ethan Breitkreutz1 个月前

@ycombinator 100%!!

Rob Thompson 的头像
Rob Thompson1 个月前

Congrats on the launch this looks sick!

James Stevens 的头像
James Stevens1 个月前

Thanks Robbie!

Mahid 的头像
Mahid1 个月前

glamorous work

Brandon Ong 的头像
Brandon Ong1 个月前

Push those 9s!

James Stevens 的头像
James Stevens1 个月前

never ending chase

Wesley Maa 的头像
Wesley Maa1 个月前

roughly how manual is this still? Like I assume enact does some kind of of OOD detection and rewinds/freezes and alerts for human takeover?

James Stevens 的头像
James Stevens1 个月前

Hey Wesley! Good q - still quite manual. When the policy does fall out of distribution, we actually don't intervene directly (yet. soon with ref: PI0.6*). We note the state(s) it failed, and then reconstruct the scene to match, and manually collect episodes there.

Aaliya 的头像
Aaliya1 个月前

congrats on the launch 99/100 rollouts after just 50 targeted recovery demos is a strong result for the DAgger approach.

James Stevens 的头像
James Stevens1 个月前

Thanks! The # of demos to get 99% will substantially rise with complexity :)

Claire Mao 的头像
Claire Mao1 个月前

this is awesome, congrats enact team!

James Stevens 的头像
James Stevens1 个月前

@clairemao78 Thanks Claire!

Ankush Dhawan 的头像
Ankush Dhawan1 个月前

Super cool

James Stevens 的头像
James Stevens1 个月前

Not as cool as you Ankush

Andre Yeung 的头像
Andre Yeung1 个月前

building next-gen agtech in stealth right now, would be interested in working with Enact to scale up our post-training/recovery data. DM sent.

Alice The Ai Expert 的头像
Alice The Ai Expert1 个月前

Enact turns real world failures into reliable robots.

Derek Askaryar 的头像
Derek Askaryar1 个月前

Incredible 😎

Teun Jansen 的头像
Teun Jansen1 个月前

Hwe have a camera like your wrist camera as well. They have a really high latency, how did you overcome this?

James Stevens 的头像
James Stevens1 个月前

the latency is fine as of now - but yeah using usb cams can be a handful

Navid Aghasadeghi 的头像
Navid Aghasadeghi1 个月前

Exactly the kind of data we need for reliable deployment. This is awesome!

Sridhar A 的头像
Sridhar A1 个月前

robots don't fail because of bad models, they fail because reality doesn't match training data. enact fixing that gap is a big deal. congrats on the launch

Jamie Ogundiran 的头像
Jamie Ogundiran1 个月前

Very cool! how do you deal with distribution shift across environments? I’d imagine the failure modes you see in your test environment can be quite different from the ones that show up in the actual deployment environment

James Stevens 的头像
James Stevens1 个月前

hey Jamie! Great question - we try to emulate production environments as much as possible to reduce that variation, i.e. try to share 90%-99% of the failure modes (not perfect).

James Stevens 的头像
James Stevens1 个月前

Note: this does constrain the prod tasks we can effectively train on, and is a very interesting economic area of exploration.

Hanming Ye 的头像
Hanming Ye1 个月前

This is going to be very useful

James Stevens 的头像
James Stevens1 个月前

the inevitable enact x waddle colab will be epic

kaan doğrusöz 的头像
kaan doğrusöz1 个月前

congrats!

James Stevens 的头像
James Stevens1 个月前

thanks Kaan!

anthony radke 的头像
anthony radke1 个月前

This looks insane! 🚀🚀🚀

Dr. Richard 的头像
Dr. Richard1 个月前

Excellent work. I encounter this exact gap in advanced manufacturing research. Post-training recovery infrastructure is what makes robotics viable beyond the lab. 👏

Mihir Rao 的头像
Mihir Rao1 个月前

Congrats James!! This is super cool

James Stevens 的头像
James Stevens1 个月前

Thanks Mihir! We need to catch up :)

Mihir Rao 的头像
Mihir Rao1 个月前

🫡

Yanda 的头像
Yanda1 个月前

Congrats on the launch!! Excited to see what DAgger at scale can do

James Stevens 的头像
James Stevens1 个月前

Will lyk as we find out!

Vladimir Karishev 的头像
Vladimir Karishev1 个月前

@lyronctk Looks super cool

Neil Nie 的头像
Neil Nie1 个月前

This is very cool! Congrats!

James Stevens 的头像
James Stevens1 个月前

thanks Neil :)

Ella Tech & Tool 的头像
Ella Tech & Tool1 个月前

Love this Targeted recovery data is exactly what real-world robotics needs

Ryan Wexler 的头像
Ryan Wexler1 个月前

@chris_j_paxton Awesome job @jamesw_stevens!

Jay@Proception 的头像
Jay@Proception1 个月前

congrats on the launch, James!

James Stevens 的头像
James Stevens1 个月前

thanks jay!

theobot 的头像
theobot1 个月前

great video!

James Stevens 的头像
James Stevens1 个月前

thanks Theo! edge of my seat for the mundane launch

Sanskar Pandey 的头像
Sanskar Pandey1 个月前

Congratulations on the launch - it’s time to hit the 3 9s!

Graham Griffin 的头像
Graham Griffin1 个月前

Let’s go James

Nicolai 的头像
Nicolai1 个月前

What arms are they? Sorry

James Stevens 的头像
James Stevens1 个月前

Hey Nicolai! YAMs. From I2rt

rohan 的头像
rohan1 个月前

nicee

James Naylor 的头像
James Naylor1 个月前

Congrats!

相关视频