正在加载视频...
视频加载失败
How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors! A fun... show more
45,721 次观看 • 18 天前 •via X (Twitter)
4 条评论

"real-time chunking (RTC)" — the quiet admission vla inference is too slow for the real world, so a small rl policy rides shotgun. if async breaks the markov assumption, how much rl theory gets rewritten just to ship robots that don't fall over?

The fresher camera input is the part that caught my attention. How does this hold up when inference latency varies, rather than staying at a fixed delay? That seems like an important test for teams running perception and control on shared hardware.

Sergey, I've followed your work for some time now and you're a heavyweight in the field of robotics. Your utilzation of integrated sensors and specialized RL environments per task is world-class!

nice framing. once perception freshness is part of the policy, inference latency stops being an infra metric and becomes a behavior variable. feels like a very underexplored lever

