Loading video...

Video Failed to Load

Go Home

Children learn from play. Can robots do the same? We propose ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ , a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with ๐‘๐€๐“๐ฌ (Robotics Agent Teams), where robots discover reusable skills through curious play. Co-led with Jiaxin Ge

95,337 views โ€ข 2 months ago โ€ขvia X (Twitter)

15 Comments

Junyi Zhang's profile picture
Junyi Zhang2 months ago

๐‘๐€๐“๐ฌ is orthogonal to current agentic robot systems. Most existing systems (like CaP-X from @letian_fu) focus on building a strong harness at test time. ๐‘๐€๐“๐ฌ runs at "play time", allowing the robot to discover skills before the task even arrives. Because they operate at different stages, skills learned by ๐‘๐€๐“๐ฌ can be directly dropped into these test-time frameworks to augment their performance.

Junyi Zhang's profile picture
Junyi Zhang2 months ago

During play, ๐‘๐€๐“๐ฌ turns open-ended exploration into reusable code skills. A team of agents repeatedly: - proposes novel-yet-learnable goals - plans with existing skills - writes robot policies as code - verifies progress step by step - diagnoses failures into feedback - stores successful behaviors in memory Play is not random: a curiosity-driven rule keeps practice at the competence frontier -- not too easy, not impossible -- and every success is distilled into a persistent code skill library.

Junyi Zhang's profile picture
Junyi Zhang2 months ago

Play pays off downstream. ๐‘๐€๐“๐ฌ improves the base CaP-Agent: LIBERO-PRO: 23.2% โ†’ 43.8% (+20.6pp) MolmoSpaces: 21.0% โ†’ 38.0% (+17.0pp) All from skills the agent acquired before it ever saw the tasks.

Junyi Zhang's profile picture
Junyi Zhang2 months ago

These play-learned skills generalize across different simulations and directly transfer to the real world. Directly using the skill library learned in LIBERO, we get: RoboSuite (cross-environment): +8.9pp Real-world tasks: +8.8pp

Junyi Zhang's profile picture
Junyi Zhang2 months ago

๐‘๐€๐“๐ฌ is a first step toward ๐๐ฅ๐š๐ฒ๐Ÿ๐ฎ๐ฅ ๐€๐ ๐ž๐ง๐ญ๐ข๐œ ๐‘๐จ๐›๐จ๐ญ ๐‹๐ž๐š๐ซ๐ง๐ข๐ง๐ : ๐ŸŒ We see a future where the next step for agentic robots isn't just stronger test-time harness, but a play stage where they set their own goals, fail, and build up skills long before we hand them a task. Huge thanks to the team: @lukehanjun (co-first) @letian_fu, Zihan Yang, Yaowei Liu, Raj Saravanan (core contributors), @istoica05 @akanazawa @JiahuiLei1998 @HavenFeng @trevordarrell and many others!

Anand Bhattad's profile picture
Anand Bhattad2 months ago

Nice work! We also have the same name: "RATS!" for Register Attention Transformers, released a few days ago, led by @TimingYang99 and @wangf3014, related to emergent part representations. More on it soon ๐Ÿ˜‰

Junyi Zhang's profile picture
Junyi Zhang2 months ago

@TimingYang99 @wangf3014 LOL, long live the RATs! ๐Ÿ€๐Ÿค

Yash's profile picture
Yash2 months ago

woah, this is interesting approach

Bilko Bibitkov's profile picture
Bilko Bibitkov2 months ago

play before the task even arrives. cool framing

Kate ๐Ÿˆ's profile picture
Kate ๐Ÿˆ2 months ago

the play-before-task stage is the part everyone skips with coding agents too. mine happily over-explores its sandbox and burns the budget before the real repo even shows up. how do you cap the play phase so it doesn't eat the whole run?

Junyi Zhang's profile picture
Junyi Zhang2 months ago

That's a nice question! Currently, we threshold the play budget so that itโ€™s comparable or less to the amount of compute used at test time. We also include a cost analysis in the appendix. That said, I think smarter ways of controlling the budget are definitely an interesting direction!

Maryam's profile picture
Maryam2 months ago

Great work!

Lรฉo's profile picture
Lรฉo2 months ago

That's a really interesting approach. It reminds me of early machine learning algorithms that tried to mimic the brain, only to raalize later this was not the way to go. But perhaps when it comes to robot behavior, mimicking a child's play time allows for efficient learning thourgh trial and error; is this where your intuition comes from?

Biketommy's profile picture
Biketommy2 months ago

Do it in human way is human level. New ways can be superintelligent level. Let AI find new ways to learn. Both known and unknown.

The AI Therapist's profile picture
The AI Therapist2 months ago

Playful learning is just reinforcement learning where the reward function is "didn't crash yet." The real question is whether the robot learns to play or just optimizes the dopamine hit. Same for us.

Related Videos

Elon just dropped a MAJOR nugget on how Tesla is going to be training Optimus to do real world tasks. They are building an Optimus Academy, which is a large scale, dedicated real-world training facility to accelerate the development of Optimus. The Academy will deploy thousands of Optimus units, potentially 10,000 to 30,000 robots, in a controlled realistic environment where they perform self-play, experiment with tasks, iterate on behaviors, and continuously generate training data through trial and error. The Tesla bots will also run millions of simulations in Teslaโ€™s high-fidelity physics-accurate engine, allowing Optimus to close the โ€œsim-to-real gapโ€ by using these real-world observations to refine and validate the simulations! โ€œYouโ€™re actually highlighting an important limitation and difference from cars. Weโ€™ll soon have 10 million cars on the road. Itโ€™s hard to duplicate that massive training flywheel. For the robot, what weโ€™re going to need to do is build a lot of robots and put them in kind of an Optimus Academy so they can do self-play in reality. Weโ€™re actually building that out. We can have at least 10,000 Optimus robots, maybe 20-30,000, that are doing self-play and testing different tasks. Tesla has quite a good reality generator, a physics-accurate reality generator, that we made for the cars. Weโ€™ll do the same thing for the robots. We actually have done that for the robots. So you have a few tens of thousands of humanoid robots doing different tasks. You can do millions of simulated robots in the simulated world. You use the tens of thousands of robots in the real world to close the simulation to reality gap. Close the sim-to-real gap.โ€

Teslaconomics

42,563 views โ€ข 7 months ago

Robotics Expert on How Robots Learn: At First Much Slower Than Humans, Then Infinitely Scalable Agility Jonathan Hurst: โ€œI don't believe that there's a singularity. I do believe that things are going to get better and better. Think of it more like a snowball picking up steam going down a hill. But the reason that it's snowballing like this is because people are putting money, and resources, and engineering time, and engineering effort in as they explore everything and start to figure all of this stuff out. Humans, for example, we've evolved to learn. We are very good at learning, and it takes very little data to show us how to do something, and then we practice and practice, iterate. Robots are not very good at learning yet. Robots take so much more data and so many more examples than a person. We're still figuring out how to teach robots how to learn. But one of the benefits that robots have in the long run is they have Wi-Fi. When you learn how to play the violin, you can't just load that to somebody else and then they know how to play the violin based on your learnings. Robots will be able to do that. @jason: โ€œOne robot learns to play violin, all robots know how to play a violin.โ€ Jonathan Hurst: โ€œOr all robots of that type know how to play the violin, right? And then minor variations for the next type and the next piece of hardware.โ€ ------------------------------ Thanks to our partners for making this possible! Most advertisers have never heard of the platform with an $11B annual run rate in ad spend. AppLovin Ads. 1B+ daily active users, full-screen video ads watched for a median of 35 seconds, and businesses are profitably spending hundreds of thousands of dollars a day on it. Now open to everyone. Sign up at If your work depends on conversations โ€” meetings, deal flow, interviews, customer calls โ€” Plaud helps you capture and organize everything with highly accurate AI-generated notes that are not just simple summaries, but also highlight pain points, key decisions, next steps, and customizable summary templates. Check out Plaud at and use code ALLIN for up to 10% off! Which is also available on Amazon: (Code: ALLIN10X)

The All-In Podcast

70,595 views โ€ข 1 month ago

New Course: ACP: Agent Communication Protocol Learn to build agents that communicate and collaborate across different frameworks using ACP in this short course built with IBM Research's BeeAI, and taught by Sandi Besen, AI Research Engineer & Ecosystem Lead at IBM, and Nicholas Renotte, Head of AI Developer Advocacy at IBM. Building a multi-agent system with agents built or used by different teams and organizations can become challenging. You may need to write custom integrations each time a team updates their agent design or changes their choice of agentic orchestration framework. The Agent Communication Protocol (ACP) is an open protocol that addresses this challenge by standardizing how agents communicate, using a unified RESTful interface that works across frameworks. In this protocol, you host an agent inside an ACP server, which handles requests from an ACP client and passes them to the appropriate agent. Using a standardized client-server interface allows multiple teams to reuse agents across projects. It also makes it easier to switch between frameworks, replace an agent with a new version, or update a multi-agent system without refactoring the entire system. In this course, youโ€™ll learn to connect agents through ACP. Youโ€™ll understand the lifecycle of an ACP Agent and how it compares to other protocols, such as MCP (Model Context Protocol) and A2A (Agent-to-Agent). Youโ€™ll build ACP-compliant agents and implement both sequential and hierarchical workflows of multiple agents collaborating using ACP. Through hands-on exercises, youโ€™ll build: - A RAG agent with CrewAI and wrap it inside an ACP server. - An ACP Client to make calls to the ACP server you created. - A sequential workflow that chains an ACP server, created with Smolagents, to the RAG agent. - A hierarchical workflow using a router agent that transforms user queries into tasks, delegated to agents available through ACP servers. - An agent that uses MCP to access tools and ACP to communicate with other agents. Youโ€™ll finish up by importing your ACP agents into the BeeAI platform, an open-source registry for discovering and sharing agents. ACP enables collaboration between agents across teams and organizations. By the end of this course, youโ€™ll be able to build ACP agents and workflows that communicate and collaborate regardless of framework. Please sign up here:

Andrew Ng

105,343 views โ€ข 1 year ago

The teams shipping AI agents right now are bleeding money on the dumbest possible expense: teaching a 400B-parameter model to read a file name. Every time an AI agent needs to "see" something today, it routes an image through a frontier model. OCR, object detection, checking if a button exists on screen. You're paying GPT-4o or Claude pricing for tasks that require perception, not reasoning. One agent workflow processing a few thousand screenshots per day can burn through more on vision calls than on the actual thinking. Perceptron's Isaac is 2B parameters. Built by the team that created Meta's Chameleon multimodal models. On perceptive benchmarks, it matches or beats models 50x its size. The VQA, OCR, and object detection scores are competitive with models running on infrastructure that costs orders of magnitude more. The MCP wrapper is the distribution play. One install command and every Claude Code agent can offload vision tasks to a model that runs on a single consumer GPU. The agent keeps its reasoning in the frontier model and routes perception to a specialist. That split is how you get vision-heavy agent workflows from "technically possible but expensive" to "cheap enough to run on everything." This is the same pattern that won in every other compute-intensive stack. General-purpose handles orchestration. Specialists handle the heavy lifting. Graphics went through it. Audio went through it. Video encoding went through it. Vision in AI agents is next. The teams building agents that see 10,000 images a day will care about this before anyone else does.

Aakash Gupta

55,978 views โ€ข 5 months ago