Loading video...

Video Failed to Load

Go Home

How can we help *any* image-input policy generalize better? 👉 Meet PEEK 🤖 — a framework that uses VLMs to decide *where* to look and *what* to do, so downstream policies — from ACT, 3D-DA, or even π₀ — generalize more effectively! 🧵

13,825 views • 11 months ago •via X (Twitter)

5 Comments

Jesse Zhang's profile picture
Jesse Zhang11 months ago

By offloading high-level reasoning of "what" and "where" to VLMs, PEEK lets robot policies focus on learning "how" to act. It’s a step toward zero-shot generalization in manipulation — giving robots a “peek” at the world.

Jesse Zhang's profile picture
Jesse Zhang11 months ago

Our policy training data consists of: - simple sim envs for 3DDA (3D and RGB input) - BRIDGE-v2 (for π₀ and ACT) This data is labeled with the PEEK VLM to directly draw “where” to focus and “what” to do onto the image.

Jesse Zhang's profile picture
Jesse Zhang11 months ago

We fine-tune policies on PEEK-VLM labeled data. We test zero-shot. 💥 41.4× sim2real improvement (3DDA) 💥 2–3.5× boosts for π₀ and ACT PEEK gives visuomotor policies the minimal cues to generalize across new tasks, scenes, and semantics.

Jesse Zhang's profile picture
Jesse Zhang11 months ago

📄 Paper: 💻 Demo: 🔗 Website, code, data, models: Work done w/ @memmelma (co-lead), @minjunkevink, @fox_dieter17849 , @_jessethomason_ , Fabio Ramos, @ebiyik_ , @abhishekunique7 , @AnqiLi24

Jesse Zhang's profile picture
Jesse Zhang11 months ago

Finally, see @memmelma's tweet thread to learn how the method works!

Related Videos

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,818 views • 1 year ago