Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

How can we help *any* image-input policy generalize better? 👉 Meet PEEK 🤖 — a framework that uses VLMs to decide *where* to look and *what* to do, so downstream policies — from ACT, 3D-DA, or even π₀ — generalize more effectively! 🧵

13,825 Aufrufe • vor 11 Monaten •via X (Twitter)

5 Kommentare

Profilbild von Jesse Zhang
Jesse Zhangvor 11 Monaten

By offloading high-level reasoning of "what" and "where" to VLMs, PEEK lets robot policies focus on learning "how" to act. It’s a step toward zero-shot generalization in manipulation — giving robots a “peek” at the world.

Profilbild von Jesse Zhang
Jesse Zhangvor 11 Monaten

Our policy training data consists of: - simple sim envs for 3DDA (3D and RGB input) - BRIDGE-v2 (for π₀ and ACT) This data is labeled with the PEEK VLM to directly draw “where” to focus and “what” to do onto the image.

Profilbild von Jesse Zhang
Jesse Zhangvor 11 Monaten

We fine-tune policies on PEEK-VLM labeled data. We test zero-shot. 💥 41.4× sim2real improvement (3DDA) 💥 2–3.5× boosts for π₀ and ACT PEEK gives visuomotor policies the minimal cues to generalize across new tasks, scenes, and semantics.

Profilbild von Jesse Zhang
Jesse Zhangvor 11 Monaten

📄 Paper: 💻 Demo: 🔗 Website, code, data, models: Work done w/ @memmelma (co-lead), @minjunkevink, @fox_dieter17849 , @_jessethomason_ , Fabio Ramos, @ebiyik_ , @abhishekunique7 , @AnqiLi24

Profilbild von Jesse Zhang
Jesse Zhangvor 11 Monaten

Finally, see @memmelma's tweet thread to learn how the method works!

Ähnliche Videos

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,818 Aufrufe • vor 1 Jahr