Загрузка видео...
Не удалось загрузить видео
VLAs? WAMs? Are language and video models the right foundations for robotics? Introducing Grounded Action Model (GAM): a new paradigm that builds robot foundation models on top of a pretrained 3D grounding model. Ground first. Then learn to act. 🧵👇
677,902 просмотров • 7 дней назад •via X (Twitter)
Комментарии: 16

1/🧵 Infants do not wait to master language before learning to reach, grasp, and explore objects. Their early interactions are grounded in perception and action. Could robot foundation models start there too: grounding objects in 3D, then learning how to act on them? "Conceptual precursors to language." by Hespos & Spelke et al supported that spatial concept comes before language in action."

2/🧵 Our idea: let pretraining provide spatial grounding, and robot demonstrations teach action. GAM builds on a pretrained 3D grounding model, in our case we used WildDet3D (you could use SAM3D or others too) to localize task objects and estimate their geometry, keeping the backbone frozen while training only the action head.

3/🧵 How does GAM act? Image tokens retain features around selected objects and the robot arm. Detection tokens encode object point clouds and metric geometry. An MM-DiT combines both with robot state history and language to generate action chunks.

4/🧵 Say it, click it, or draw a box. GAM combines language for what to do with points or boxes for which objects to act on. This enables more grounded action specification, reduces language ambiguity, and gives high-level VLMs a direct interface to guide GAM.

5/🧵 Move the object. Select a different target. The grounded observation follows that selection and updates its geometry. The goal is not to ignore visual changes. It is to respond to task-relevant spatial changes while reducing exposure to irrelevant scene variation.

6/🧵 RoboTwin 2.0, 50 tasks: 55.3% average success vs. 52.0% for Spatial Forcing. 47.6% on randomized scenes vs. 30.4% for Abot-M0. Action policies train per task on clean demos. The grounding backbone is adapted to simulator images separately, then frozen.

7/🧵 LIBERO-PRO: 61% average success across 16 settings, vs. 53% for π0.5. When the target changes in LIBERO-Spatial, GAM reaches 88% vs. 1% for π0.5 and much better than Flexi-π another recent 3D WAM. The objects and actions are familiar; the challenge is acting on the newly selected target.

8/🧵 For long-horizon and memory-dependent tasks, a Molmo2 planner uses current observations and episode history to select targets via points. GAM handles the actions. On Franka, GAM + Molmo2 achieves 64.7% ID and 49.8% OOD step completion.

9/🧵 GAM is a starting point. Grounding errors and omitted context, such as unselected obstacles, remain challenges. We see an opportunity to scale 3D grounding as a foundation for robot learning, alongside language and planning. Paper: Code: Project:

This project is led by @gehao_zhang_ and collaborated with @weikaih04 @Shailes_h_ @RanjayKrishna And @gehao_zhang_ is amazing, and will be looking for a PhD this coming cycle! Definitely an amazing candidate!

wow!

Explicit 3D grounding could make failures easier to diagnose: did the robot locate the wrong object, estimate its geometry poorly, or choose the wrong action? I'd be interested in that breakdown alongside success rates. It could tell a small team what data to collect next.

ground first, act second. makes a lot of sense for real-world robotics

damn, this is actually a really smart approach.

Space before skill.

47.6% on randomized scenes is the number that matters most. If grounding transfers while language priors collapse under shift, the real test is objects never seen in grounding pretraining. How far does zero-shot hold when the target object class was never grounded?
