Loading video...

Video Failed to Load

Go Home

How does test-time scaling impact robots? We find that larger models, more thinking, and more context help significantly for some prompts but not others. Like LLMs, we can also train a router to for a better performance/latency tradeoff! Paper:

25,235 views • 3 months ago •via X (Twitter)

3 Comments

Mingkai Deng's profile picture
Mingkai Deng3 months ago

Really cool results! We've been studying the same question using agentic LLMs for demonstration. Here's the twist: instead of an external router over a hand-enumerated planner pool, we built self-regulation as the model's own decision, so it's optimized end-to-end with RL. This way, the regulation strategies *emerge* rather than being picked from a menu. After RL, our model learned to plan further ahead per invocation (+22.8% horizon) while barely planning more often (+2%)

Stian Jakobsen's profile picture
Stian Jakobsen3 months ago

interesting it helps on some prompts but not others. feels like the gains land when the task is reasoning-bound, not dynamics-bound. does that track?

GEE!'s profile picture
GEE!3 months ago

Makes sense, the 'think before you act' principle applies to robots too.

Related Videos