Loading video...
Video Failed to Load
How does test-time scaling impact robots? We find that larger models, more thinking, and more context help significantly for some prompts but not others. Like LLMs, we can also train a router to for a better performance/latency tradeoff! Paper:
25,235 views • 3 months ago •via X (Twitter)
3 Comments

Really cool results! We've been studying the same question using agentic LLMs for demonstration. Here's the twist: instead of an external router over a hand-enumerated planner pool, we built self-regulation as the model's own decision, so it's optimized end-to-end with RL. This way, the regulation strategies *emerge* rather than being picked from a menu. After RL, our model learned to plan further ahead per invocation (+22.8% horizon) while barely planning more often (+2%)

interesting it helps on some prompts but not others. feels like the gains land when the task is reasoning-bound, not dynamics-bound. does that track?

Makes sense, the 'think before you act' principle applies to robots too.

