Video wird geladen...
Video konnte nicht geladen werden
๐ Introducing LeVERB, the first ๐น๐ฎ๐๐ฒ๐ป๐ ๐๐ต๐ผ๐น๐ฒ-๐ฏ๐ผ๐ฑ๐ ๐ต๐๐บ๐ฎ๐ป๐ผ๐ถ๐ฑ ๐ฉ๐๐ (upper- & lower-body), trained on sim data and zero-shot deployed. Addressing interactive tasks: navigation, sitting, locomotion with verbal instruction. ๐งต
96,617 Aufrufe โข vor 1 Jahr โขvia X (Twitter)
11 Kommentare

(1/9) Old approaches made humanoid robots follow hand-crafted action commands (like setting a walking speed or an arm pose) from the language module. This limited them to a small, predefined skill set and made complex whole-body motions hard to achieve.

(2/9) LeVERB instead learns a latent action space, shout out to PULSE and MaskedMimic. The high-level VL model outputs a latent โverbโ, the low-level controller decodes it into joint motion (seperatedly trained). Result: a far richer expressive skill set.

(3/9) Two-system brain: System 2 thinks at 10 Hz (vision + language). System 1 reacts at 50 Hz (balance + contacts). Slow reasoning + fast reflexes = stable, expressive whole-body control.

(4/9) Training data is the real bottleneck, so we built LeVERB-Bench: 154+ photorealistic sim scenes with heavy randomizationโlighting, textures, clutter, camera angles. This diversity is what lets LeVERB generalize. Tasks: visual navigation, sitting, reaching, locomotion, etc.

(5/9) LeVERB sees 80% zero-shot success rate on simple visual navigation tasks, and 58.5% across the board. This is 7.8 times better than a naive hierarchical VLA implementation with no latent regularization, highlighting the unique challenge brought by async decoupled WBC loop.

(6/9) Generalization: LeVERB sees โtake a seatโ, โsit downโ, or โsit on blue chairโ and knows they mean the same thing. It also reasons about space: if the chair is in front, it turns first, then sits.

(7/9) Related inspiration: Helix, NaVILA, LangWBC, VBC, โฆ each pushes latent or hierarchical VLA in its own niche (upper-body, legged nav, etc.). LeVERB adds whole body latent control plus an open benchmark for everyone to build on.

(8/9) Current status: dynamics-level sim2real โ๏ธ vision sim2real - to be released โข LeVERB-Bench dataset is already open-sourced in LeRobot format โข Full code release is coming

(9/9) Dive deeper ๐ Collaborators: @x_h_ucb @Dantong_Niu @qiayuanliao @tjomiii Jan Tommy Gravdahl @xbpeng4 @GuanyaShi @trevordarrell @KoushilSreenath Shankar Sastry

so cool! integrated in @LeRobotHF ?

@LeRobotHF only the dataset so far haha would like to see if model can be integrated too ๐
