Video yükleniyor...
Video Yüklenemedi
Introducing ACT-2 Preview The first robotics model to unify broad generalization with high reliability. A single fine-tuning example can teach Memo a new behavior that generalizes. Zero shot, real unseen homes, 99% success rate.
1,042,998 görüntüleme • 1 ay önce •via X (Twitter)
61 Yorum

We completed the most rigorous generalization test to date. Across 785 trials in 31 unseen environments, ACT-2 achieved 99.1% success in laundry folding. ACT-2 also achieved human-level fold quality: receiving an average rating of 4.72/5, with 98.3% earning four or five stars.

Tons of emergent behavior along the way: - Picking up clothes off the ground - Handling baby wear to 8XL shirts - Robustness against adversarial disturbances and lighting

We call this a Solve. Progress in robotics is difficult to measure because demos vary by setup. Demo ≠ Solved. A Solve declares two boundaries: scope and adaptation cost. Without both, 99% has no context.

We found a general recipe for Solves: scale pretraining, then hill-climb with minimal in-house data. For the first time, one fine-tuning example can teach a new behavior that generalizes. Below: 4 folding strategies, learned from one example each and tested on held-out setups.

Quantitatively, we measure the generalization gap as the difference between in-domain and out-of-domain performance. As we scale up pretraining, the gap falls sharply. This makes in-house performance a reliable predictor of performance in the wild.

This property allows us to hill-climb performance in our office, and trust those gains to hold in unseen homes Our fleet of Memos runs in parallel to rapidly advance reliability, quality, and speed. Left: fleet-scale improvement in-house Right: Memo working across unseen homes

Laundry is our first Solve of many. Our recipe is so general that scaling data and compute gives us predictable improvements. Unlocking one Solve accelerates the next Solve. The same ACT-2 model is learning to vacuum, organize toys, zip clothing, and turn pants inside out.

This fall, ACT-2’s first Solve enters homes through our Beta Program, the final step towards fully autonomous home robot deployment. Full technical report:

So cool, can't wait to have one at my place

Very soon 😉

Is it possible to get access to the pretrained model? We recently developed a mechanistic interpretability concept for fast-adaptation by tuning only the relevant parts of the network: If the pretrained model is strong, we can do magic. Give us access, and we can unlock continual learning together ;)

This is very cool, checking it out now!

sending this to my mom

Very cool to see this kind of generalization @tonyzzhao

Just ask for generalization! @ericjang11

But are you still shipping this year

Have you applied!

@timzaman Congrats @tonyzzhao on the launch!!! Can’t wait to have one at home.

@timzaman Thank you Yuanhao! Apply here!

Congrats! data engine end to end ftw 💪

Extremely cool ;)

Thank you Remi!! Congrats on the humanoid release. So cool.

congrats tony! awesome to see generalization + reliability

Thank you Jason! 🙏

The wanting intensifies! Does this mean I will be able to teach it to fold things my way? 😍

Exactly! I'm honestly quite surprised that it just works with SFT..!

Proud of the team for this achievement

@RemiCadene Very Impressive. Congratulations! There is a snapshot where the head camera is blocked, and the robot can still keep going. How to explain this behavior?

@RemiCadene ACT-2 takes all 5 cameras (one head, two on each hand) as input. And sometimes dropout is all you need!

Congratulations @sundayrobotics team! Incredible work

@sundayrobotics Thank you Ryan!

@Scobleizer I can’t believe nobody has asked for the holy grail of folding: The King Size Fitted Sheet.

big breakthrough - congrats!

Thank you Aaref for being part of this journey ❤️

This all seemed so far away till last year

Really cool work! It will be cooler if the weights could be open source, seems to be a common question here.

Congrats @tonyzzhao & team!

Thank you Lindon!!

Generalizing reliability!!

amazing, can you share the absolute number of episodes in the pre-training dataset?

These results are pretty impressive. So this is what the Sunday team meant when they said most robotics companies were collecting data the wrong way.

@RemiCadene Congrats, very exciting!

@RemiCadene Thank you Ted! 🙏

Impressive, but let's see it do all laundry, not just folding. Folding doesn't really change depending on where you are, full laundry definitely does

This is cool, definitely one of the hard problems

It was very difficult. We weren’t sure if it’s possible beginning of the year.

We all know what Green cap in Chinese culture means…

@philfung What about the emergent behavior you said you saw? That’s amazing but you left us curious!

Congrats! Looking solid!

Thank you Haoru! Precise posting after vague posting 😉

How do you think about a "solve" when the cost of errors scales? Like 99% in self driving is unlaunchable.

Precisely. The performance threshold could differ across applications. What a Solve highlights is that we should not omit Scope and Adaptation Budget when reporting the success rate. The Solve framework does not carry any opinion about the performance threshold itself.

Very cool congrats!

Thank you Shuang!

Very cool videos! I appreciate the detailed descriptions of the evals in the blog. Is there a model card somewhere?

Thank you Dhruv. We don't have it right now but might release it in the future with the full release 🫡

Look forward to reading it if/when you do. Also signed up for the waitlist — kudos on the launch!

so cool, congrats

@wenlong_huang @tonyzzhao - it’s incredible how far ACT is going! Let there be no boundaries. Kudos to entire @sundayrobotics team

@arthurallshire Congrats Tony, really cool results!

@arthurallshire Thank you Zeeshan! Hope all is well and congrats on the new journey!
