Загрузка видео...
Не удалось загрузить видео
Today we’re launching Enact: post-training infrastructure that makes robotics models work in the real world. Robotics models fail when execution reaches states their training data never covered. A slip, an off-angle grasp, or an external disturbance can leave the policy without a learned recovery. Enact finds those failures and... show more
52,909 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 62

Exciting release! We need more data groups that deeply understand DAgger

Thanks Kyle :)

To isolate why recovery data matters, we used a controlled packing task common in e-commerce fulfillment, with a single item. We collected 200 demonstrations of successful placements into the box (the “happy path”) and fine-tuned π0.5. The policy failed on 10 of 100 rollouts. Adding more happy-path demonstrations did not improve success rates. Targeted recovery data did.

Why? Happy-path demonstrations do not cover the out-of-distribution states created by the policy’s own mistakes. Every example showed a successful handoff, so the policy received no training signal for recovering after a drop.

After adding 50 targeted recovery demonstrations, the retrained policy succeeded on 99 of 100 consecutive packing rollouts. It recovered from every table drop.

This is a simple application of Dataset Aggregation (DAgger): roll out the current policy, collect targeted recovery demonstrations at the failure states it actually reaches, aggregate them into the training set, and retrain. As task complexity increases, failure modes multiply, rare failures become harder to discover, and collecting the recovery data needed to solve them takes substantially more effort.

In our table-bussing task, the obvious failures were cups or plates flipping. Those were easy to predict and collect recovery data for. But roughly once every 100 runs, the knife wedged itself under the bin.

The only way to recover was to move the bin itself, which meant training the robot on an entirely new skill! This is the long tail that breaks deployed systems: a rare failure can demand a behavior the original task never required. Finding these states and generating the data to solve them is what Enact does.

We’re already serving our first customers and applying this loop to failures their models encounter in the real world. Deploying a policy that fails on a real task? Tell us the model, task, and where it breaks: [email protected] or DM me.

This is awesome

@ycombinator This is really interesting. Excited to see the progression as you guys continue forward.

@ethan_breitk @ycombinator Thanks Ethan! Stay tuned ;)

@ycombinator 100%!!

Congrats on the launch this looks sick!

Thanks Robbie!

glamorous work

Push those 9s!

never ending chase

roughly how manual is this still? Like I assume enact does some kind of of OOD detection and rewinds/freezes and alerts for human takeover?

Hey Wesley! Good q - still quite manual. When the policy does fall out of distribution, we actually don't intervene directly (yet. soon with ref: PI0.6*). We note the state(s) it failed, and then reconstruct the scene to match, and manually collect episodes there.

congrats on the launch 99/100 rollouts after just 50 targeted recovery demos is a strong result for the DAgger approach.

Thanks! The # of demos to get 99% will substantially rise with complexity :)

this is awesome, congrats enact team!

@clairemao78 Thanks Claire!

Super cool

Not as cool as you Ankush

building next-gen agtech in stealth right now, would be interested in working with Enact to scale up our post-training/recovery data. DM sent.

Enact turns real world failures into reliable robots.

Incredible 😎

Hwe have a camera like your wrist camera as well. They have a really high latency, how did you overcome this?

the latency is fine as of now - but yeah using usb cams can be a handful

Exactly the kind of data we need for reliable deployment. This is awesome!

robots don't fail because of bad models, they fail because reality doesn't match training data. enact fixing that gap is a big deal. congrats on the launch

Very cool! how do you deal with distribution shift across environments? I’d imagine the failure modes you see in your test environment can be quite different from the ones that show up in the actual deployment environment

hey Jamie! Great question - we try to emulate production environments as much as possible to reduce that variation, i.e. try to share 90%-99% of the failure modes (not perfect).

Note: this does constrain the prod tasks we can effectively train on, and is a very interesting economic area of exploration.

This is going to be very useful

the inevitable enact x waddle colab will be epic

congrats!

thanks Kaan!

This looks insane! 🚀🚀🚀

Excellent work. I encounter this exact gap in advanced manufacturing research. Post-training recovery infrastructure is what makes robotics viable beyond the lab. 👏

Congrats James!! This is super cool

Thanks Mihir! We need to catch up :)

🫡

Congrats on the launch!! Excited to see what DAgger at scale can do

Will lyk as we find out!

@lyronctk Looks super cool

This is very cool! Congrats!

thanks Neil :)

Love this Targeted recovery data is exactly what real-world robotics needs

@chris_j_paxton Awesome job @jamesw_stevens!

congrats on the launch, James!

thanks jay!

great video!

thanks Theo! edge of my seat for the mundane launch

Congratulations on the launch - it’s time to hit the 3 9s!

Let’s go James

What arms are they? Sorry

Hey Nicolai! YAMs. From I2rt

nicee

Congrats!

