Загрузка видео...

Не удалось загрузить видео

На главную

VLAs? WAMs? Are language and video models the right foundations for robotics? Introducing Grounded Action Model (GAM): a new paradigm that builds robot foundation models on top of a pretrained 3D grounding model. Ground first. Then learn to act. 🧵👇

677,902 просмотров • 7 дней назад •via X (Twitter)

Комментарии: 16

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

1/🧵 Infants do not wait to master language before learning to reach, grasp, and explore objects. Their early interactions are grounded in perception and action. Could robot foundation models start there too: grounding objects in 3D, then learning how to act on them? "Conceptual precursors to language." by Hespos & Spelke et al supported that spatial concept comes before language in action."

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

2/🧵 Our idea: let pretraining provide spatial grounding, and robot demonstrations teach action. GAM builds on a pretrained 3D grounding model, in our case we used WildDet3D (you could use SAM3D or others too) to localize task objects and estimate their geometry, keeping the backbone frozen while training only the action head.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

3/🧵 How does GAM act? Image tokens retain features around selected objects and the robot arm. Detection tokens encode object point clouds and metric geometry. An MM-DiT combines both with robot state history and language to generate action chunks.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

4/🧵 Say it, click it, or draw a box. GAM combines language for what to do with points or boxes for which objects to act on. This enables more grounded action specification, reduces language ambiguity, and gives high-level VLMs a direct interface to guide GAM.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

5/🧵 Move the object. Select a different target. The grounded observation follows that selection and updates its geometry. The goal is not to ignore visual changes. It is to respond to task-relevant spatial changes while reducing exposure to irrelevant scene variation.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

6/🧵 RoboTwin 2.0, 50 tasks: 55.3% average success vs. 52.0% for Spatial Forcing. 47.6% on randomized scenes vs. 30.4% for Abot-M0. Action policies train per task on clean demos. The grounding backbone is adapted to simulator images separately, then frozen.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

7/🧵 LIBERO-PRO: 61% average success across 16 settings, vs. 53% for π0.5. When the target changes in LIBERO-Spatial, GAM reaches 88% vs. 1% for π0.5 and much better than Flexi-π another recent 3D WAM. The objects and actions are familiar; the challenge is acting on the newly selected target.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

8/🧵 For long-horizon and memory-dependent tasks, a Molmo2 planner uses current observations and episode history to select targets via points. GAM handles the actions. On Franka, GAM + Molmo2 achieves 64.7% ID and 49.8% OOD step completion.

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

9/🧵 GAM is a starting point. Grounding errors and omitted context, such as unselected obstacles, remain challenges. We see an opportunity to scale 3D grounding as a foundation for robot learning, alongside language and planning. Paper: Code: Project:

Фото профиля Jiafei Duan
Jiafei Duan7 дней назад

This project is led by @gehao_zhang_ and collaborated with @weikaih04 @Shailes_h_ @RanjayKrishna And @gehao_zhang_ is amazing, and will be looking for a PhD this coming cycle! Definitely an amazing candidate!

Фото профиля atharva ☆
atharva ☆7 дней назад

wow!

Фото профиля Abby
Abby7 дней назад

Explicit 3D grounding could make failures easier to diagnose: did the robot locate the wrong object, estimate its geometry poorly, or choose the wrong action? I'd be interested in that breakdown alongside success rates. It could tell a small team what data to collect next.

Фото профиля SEAR
SEAR7 дней назад

ground first, act second. makes a lot of sense for real-world robotics

Фото профиля Pawan
Pawan6 дней назад

damn, this is actually a really smart approach.

Фото профиля Human
Human6 дней назад

Space before skill.

Фото профиля XoroAI
XoroAI7 дней назад

47.6% on randomized scenes is the number that matters most. If grounding transfers while language priors collapse under shift, the real test is objects never seen in grounding pretraining. How far does zero-shot hold when the target object class was never grounded?

Похожие видео

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 просмотров • 3 лет назад