Video wird geladen...
Video konnte nicht geladen werden
Practice makes perfect. A truly agentic robot shouldn’t just know what to do. It should also figure out how to get better. What if a robot could watch a demonstration, figure out what it needs to practice, and learn how to get better through practice? Introducing RPG: Guided Self-Improvement... show more
160,963 Aufrufe • vor 2 Tagen •via X (Twitter)
8 Kommentare

What should the robot practice? RPG uses real-world demonstrations to reconstruct simplified simulation tasks that target the capabilities needed for deployment. These tasks need not exactly reproduce the real-world scene. They need to provide useful practice for learning the required skills.

How should the robot improve through practice? RPG uses execution feedback, privileged simulator state, and a video analyzer to diagnose failures. Demonstrations provide a reference for guiding improvement. The agent adds and revises reusable skills, then tests the changes across tasks before retaining them in a shared skill library.

15 rounds of practice. 28.6% → 95.0%. Across 22 manipulation tasks: • RPG (Gemini 3.8 Flash): 95.0% • ASPIRE: 75.5% • GPT-6 Astra Pro + CaP-Agent0: 60.0% • Gemini + CaP-Agent0: 48.6% • RATs: 41.8% RPG gets there by improving the execution system and shared skills — without updating the foundation model weights. The model stays the same. The agent system gets better.

The real test is whether the improvements learned in simulation transfer back to the real world. After practice, we freeze the improved system and deploy it on the physical robot. RPG achieves 30/30 successful real-world trials across: 🗄️ Store ball in drawer and close it 🧣 towel folding 🥣 Transfer bowl from left to right The improved system also transfers to additional real-world manipulation tasks. Practice in sim. Improve the system. Go real.

So what actually changes as RPG practices? Over 15 rounds, the shared skill library grows from 15 → 38 skills, with: ➕ 23 new skills 🔧 66 revisions Failures become feedback. Feedback becomes reusable skills. Reusable skills help future tasks. Our takeaway: You don’t need a perfect digital twin to get useful sim-to-real improvement. Even with simplified reconstructed environments and practice tasks, coding agents can learn execution improvements that transfer surprisingly well to the real world. Huge thanks to all of our collaborators and co-authors who made this possible — @erichzjiang, @DengShuying, @HaoruXue, @Weirui_Ye, @rocky_duan, @nhaghtal, Shankar Sastry, @pabbeel, and @HaozhiQ! 🌐 Project: 📄 Paper: Also checkout concurrent work SimEX ( from friends in Amazon FAR!

robot practices for 15 rounds and still beats my benchmark scores smh

🤣

Picking what to practice is the underrated half. Retrying is cheap for any agent. Knowing which failure is worth a hundred more attempts, and which one was noise, is where the learning actually comes from.
