Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

How to harness foundation models for *generalization in the wild* in robot manipulation? Introducing VoxPoser: use LLM+VLM to label affordances and constraints directly in 3D perceptual space for zero-shot robot manipulation in the real world! 🌐 🧵👇

293,910 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

Data is key for generalization, but robot data is scarce and expensive. Instead of training policies on labeled data, VoxPoser uses LLM+VLM to compose 3D value maps using generated code. Then 6-DoF actions are synthesized by motion planners, all w/o any training or primitives.

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

Given free-form instructions + RGB-D obs, LLM orchestrates perception calls to VLM and array operations to assign continuous values to voxel map, showing *where to act* and *how to act*. It also parametrizes rotation, velocity, and gripper actions for a complete SE(3) trajectory!

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

We verified VoxPoser in everyday manipulation tasks in the wild, including articulated and deformable object manipulation. All the results here are synthesized with zero-shot execution.

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

Just toss your objects too! VoxPoser is robust to disturbances because it replans actions in *real-time* with visual feedback. The 3D value maps are always updated with latest observations, allowing robot to recover from unexpected errors.

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

LLMs show emergent abilities at scale – same applies to VoxPoser, but on physical behaviors! It can conduct physics experiments, have behavioral commonsense, listen to your fine-grained correction, come up with multi-step visual program, and more.

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

For more, check out 👇 Project site: Walkthrough video: Paper: Work done w/ @chenwang_j , @RuohanZhang76 , @YunzhuLiYZ , @jiajunwu_cs, and @drfeifei at @StanfordSVL & @StanfordAILab.

Profilbild von Wenlong Huang
Wenlong Huangvor 2 Jahren

@RuohanZhang76 @YunzhuLiYZ @jiajunwu_cs @drfeifei @StanfordSVL @StanfordAILab Excited to share that we have open-sourced the code based on RLBench: We will also be presenting VoxPoser at CoRL next week as an oral presentation. See you in Atlanta! #CoRL @corl_conf

Profilbild von Brett Adcock
Brett Adcockvor 3 Jahren

This is great

Profilbild von Wenlong Huang
Wenlong Huangvor 3 Jahren

Thank you Brett!

Profilbild von DK Xu
DK Xuvor 3 Jahren

interesting work!!

Ähnliche Videos