正在加载视频...
视频加载失败
🧐Applying world models to improve real-world policy on challenging manipulation tasks used to be considered out of reach. 😌After sustained effort, we’re now seeing encouraging progress. 🚀Thrilled to introduce RISE: Self-Improving Robot Policy with Compositional World Model RISE is, to our knowledge, the first work to use a world... show more
77,124 次观看 • 7 个月前 •via X (Twitter)
24 条评论

RISE (1/N) 🤖Before diving into technical details, let's first see what we've achieved on challenging manipulation. RISE yields absolute performance increases of: +45% in Backpack Packing +35% in Box Closing +35% in Dynamic Brick Sorting compared to a strong offline RL method (RECAP, re-implementated on pi_0.5, which is an amazing work by @physical_int) These improvements highlight the effectiveness of leveraging a proper world model for policy optimization for contact-rich manipulation.

RISE (2/N) 🥲A central challenge to (sensory) robotic reinforcement learning is the restriction of physical interactions, which is unparalleled and costly in both hardware and human labor. Consequently, many real-world robotic RL pipelines rely heavily on offline datasets and online interactions from a small number of deployed robots, which constrain the overall policy improvement.

RISE (3/N) To address this bottleneck, we introduce RISE: Reinforcement learning via Imagination for SElf-improving robots. RISE shifts the learning environment from physical world to a Compositional World Model, which first emulates future observations for proposed actions, then evaluates imagined states to derive advantage for policy improvement.

RISE (4/N) One key contribution is an interactive yet imaginative learning environment enabled by a Compositional World Model that factorizes world modeling into two complementary components: - A dynamics model for future-state generation under candidate actions; - A value model for state evaluation, which further forms a dense advantage for policy optimization. This decomposition enables architectural and objective specialization for each sub-module.

RISE (5/N) Upon the learned world model, the self-improving loop for policy optimization encompasses two stages. 1) Rollout stage. Prompted with an optimal advantage, the rollout policy interacts with the world model to produce rollout data. 2) Training stage. The behavior policy is then trained to generate proper action under an advantage-conditioning scheme.

RISE (6/N) Main results on three tasks.

RISE (7/N) 🥳For more details, please refer to our project page: and paper:

RISE (8/N) 💖Importantly, RISE is built upon the collective effort of the growing robot learning community, and we want to share and pay tribute to the following pioneers and projects: - @JasonMa2020: large-scale value learning upon vision-language models. E.g., GVL: - @physical_int @chelseabfinn: foundation policy and advantage-conditioned RL formulation, enabling IL-style optimization despite the complexity of action denoising. I.E., - @du_yilun: series of studies on developing highly capable systems via carefully composing heterogeneous foundation models. E.g., - @LTXStudio @ltx_model: for delivering highly efficient video models, which are important to be integrated into the RL loop. - @AGIBOTofficial @GalaxeaDynamics: for releasing large-scale robot datasets, where the dynamics model is initialized from. -@DrJimFan: for always giving hope on leveraging video representation for robot learning. See also their recent work: DreamZero by @jang_yoel @ShenyuanGao

@hendrikomg @isaacery1

Nice. Just emailed you.

Cool work! 🦾

Thank you!!

Struggling with robotics simulations? 🤖 We are a team of ex-Amazon Roboticist and launched a free, open-source AI tool to build simulations for you! Check it out: We are giving away a Claude Code to 5 people who try it and share feedback via a quick interview, DM me for infos.

Congrats on the work! Is the world model online-updated with the real world rollouts?

Thank you Chenhao. In this version, the world model is offline built and kept frozen during RL stage. We will try ti make it adaptive to new streaming data.

Thanks for the explanation! If this is the case, how would it compared with training policies completely offline in the world model?

Great work! It’s interesting that you guys say recap only achieves 60-40% success rate. Is this only for one rollout stage?

Thank you for your interest! Yes, one rollout stage only. One major reason for the moderate recap performance is the policy pretraining. In Pi06*, recap method is applied in both policy pretraining and finetuning, upon a proprietary Pi06 architecture, whereas we can only re-implement it on a public open-pi05 checkpoint, which is not pretrained with action conditioning as additional input.

Ah I see, makes sense! Thank you for the clarification

Truly admire the work you have done with world model, Jiazhi. May I have the honor to learn more from you?

Great work

Thank you!!

Groundbreaking leap for robotic autonomy!

Thank you for your kind words!
