Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Practice makes perfect. A truly agentic robot shouldn’t just know what to do. It should also figure out how to get better. What if a robot could watch a demonstration, figure out what it needs to practice, and learn how to get better through practice? Introducing RPG: Guided Self-Improvement...

160,963 Aufrufe • vor 2 Tagen •via X (Twitter)

8 Kommentare

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

What should the robot practice? RPG uses real-world demonstrations to reconstruct simplified simulation tasks that target the capabilities needed for deployment. These tasks need not exactly reproduce the real-world scene. They need to provide useful practice for learning the required skills.

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

How should the robot improve through practice? RPG uses execution feedback, privileged simulator state, and a video analyzer to diagnose failures. Demonstrations provide a reference for guiding improvement. The agent adds and revises reusable skills, then tests the changes across tasks before retaining them in a shared skill library.

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

15 rounds of practice. 28.6% → 95.0%. Across 22 manipulation tasks: • RPG (Gemini 3.8 Flash): 95.0% • ASPIRE: 75.5% • GPT-6 Astra Pro + CaP-Agent0: 60.0% • Gemini + CaP-Agent0: 48.6% • RATs: 41.8% RPG gets there by improving the execution system and shared skills — without updating the foundation model weights. The model stays the same. The agent system gets better.

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

The real test is whether the improvements learned in simulation transfer back to the real world. After practice, we freeze the improved system and deploy it on the physical robot. RPG achieves 30/30 successful real-world trials across: 🗄️ Store ball in drawer and close it 🧣 towel folding 🥣 Transfer bowl from left to right The improved system also transfers to additional real-world manipulation tasks. Practice in sim. Improve the system. Go real.

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

So what actually changes as RPG practices? Over 15 rounds, the shared skill library grows from 15 → 38 skills, with: ➕ 23 new skills 🔧 66 revisions Failures become feedback. Feedback becomes reusable skills. Reusable skills help future tasks. Our takeaway: You don’t need a perfect digital twin to get useful sim-to-real improvement. Even with simplified reconstructed environments and practice tasks, coding agents can learn execution improvements that transfer surprisingly well to the real world. Huge thanks to all of our collaborators and co-authors who made this possible — @erichzjiang, @DengShuying, @HaoruXue, @Weirui_Ye, @rocky_duan, @nhaghtal, Shankar Sastry, @pabbeel, and @HaozhiQ! 🌐 Project: 📄 Paper: Also checkout concurrent work SimEX ( from friends in Amazon FAR!

Profilbild von Hershal Rao
Hershal Raovor 2 Tagen

robot practices for 15 rounds and still beats my benchmark scores smh

Profilbild von Yen-Jen Wang
Yen-Jen Wangvor 2 Tagen

🤣

Profilbild von Karan Jagtiani
Karan Jagtianivor 2 Tagen

Picking what to practice is the underrated half. Retrying is cheap for any agent. Knowing which failure is worth a hundred more attempts, and which one was noise, is where the learning actually comes from.

Ähnliche Videos

Elon just dropped a MAJOR nugget on how Tesla is going to be training Optimus to do real world tasks. They are building an Optimus Academy, which is a large scale, dedicated real-world training facility to accelerate the development of Optimus. The Academy will deploy thousands of Optimus units, potentially 10,000 to 30,000 robots, in a controlled realistic environment where they perform self-play, experiment with tasks, iterate on behaviors, and continuously generate training data through trial and error. The Tesla bots will also run millions of simulations in Tesla’s high-fidelity physics-accurate engine, allowing Optimus to close the “sim-to-real gap” by using these real-world observations to refine and validate the simulations! “You’re actually highlighting an important limitation and difference from cars. We’ll soon have 10 million cars on the road. It’s hard to duplicate that massive training flywheel. For the robot, what we’re going to need to do is build a lot of robots and put them in kind of an Optimus Academy so they can do self-play in reality. We’re actually building that out. We can have at least 10,000 Optimus robots, maybe 20-30,000, that are doing self-play and testing different tasks. Tesla has quite a good reality generator, a physics-accurate reality generator, that we made for the cars. We’ll do the same thing for the robots. We actually have done that for the robots. So you have a few tens of thousands of humanoid robots doing different tasks. You can do millions of simulated robots in the simulated world. You use the tens of thousands of robots in the real world to close the simulation to reality gap. Close the sim-to-real gap.”

Teslaconomics

42,563 Aufrufe • vor 8 Monaten

Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library. ASPIRE is a new type of continual learning: "training" is skill refinement instead of gradient descent. "Trained model" is a repo of sensorimotor skills instead of floating weights. “Distributed training” is a panel of agents each practicing a different skill instead of sharded minibatches. Here's the beauty: ASPIRE gives the tired terms "sim2real transfer" and "cross-embodiment transfer" a whole new meaning. Bridging the sim-to-real gap is notoriously brutal. An end-to-end policy has to swallow both the visual shift (sim looks toyish next to a real camera) and the subtle contact physics it never quite gets right. ASPIRE sidesteps the mess, because it doesn't ship pixels or weights across the gap, but ships the know-how. The robot still has to practice in the real world, not zero-shot, but it gets there way faster because it isn't rediscovering the strategy from scratch. Same for going single-arm to bimanual hardware, which usually requires new data and retraining from zero. ASPIRE achieves up to ~10x cut in "transfer learning” tokens (yes, tokens are the new unit of *training* compute ;) Check out our gallery of 150+ tasks and 90+ skills the robots taught themselves, all on the website! Kind of wild that we can ship the "learned weights" as an HTML page rather than a GGUF. We'll open-source the full stack so your own robot library starts compounding from ours! Deep dive in thread:

Jim Fan

215,578 Aufrufe • vor 3 Monaten