正在加载视频...

视频加载失败

How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors! A fun...

45,721 次观看 • 18 天前 •via X (Twitter)

4 条评论

Stephen 的头像
Stephen18 天前

"real-time chunking (RTC)" — the quiet admission vla inference is too slow for the real world, so a small rl policy rides shotgun. if async breaks the markov assumption, how much rl theory gets rewritten just to ship robots that don't fall over?

Abby 的头像
Abby17 天前

The fresher camera input is the part that caught my attention. How does this hold up when inference latency varies, rather than staying at a fixed delay? That seems like an important test for teams running perception and control on shared hardware.

JK 的头像
JK18 天前

Sergey, I've followed your work for some time now and you're a heavyweight in the field of robotics. Your utilzation of integrated sensors and specialized RL environments per task is world-class!

Noorullah Jamakzai 的头像
Noorullah Jamakzai17 天前

nice framing. once perception freshness is part of the policy, inference latency stops being an infra metric and becomes a behavior variable. feels like a very underexplored lever

相关视频

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 次观看 • 7 个月前