Video yükleniyor...
Video Yüklenemedi
How can we help *any* image-input policy generalize better? 👉 Meet PEEK 🤖 — a framework that uses VLMs to decide *where* to look and *what* to do, so downstream policies — from ACT, 3D-DA, or even π₀ — generalize more effectively! 🧵
13,825 görüntüleme • 11 ay önce •via X (Twitter)
5 Yorum

By offloading high-level reasoning of "what" and "where" to VLMs, PEEK lets robot policies focus on learning "how" to act. It’s a step toward zero-shot generalization in manipulation — giving robots a “peek” at the world.

Our policy training data consists of: - simple sim envs for 3DDA (3D and RGB input) - BRIDGE-v2 (for π₀ and ACT) This data is labeled with the PEEK VLM to directly draw “where” to focus and “what” to do onto the image.

We fine-tune policies on PEEK-VLM labeled data. We test zero-shot. 💥 41.4× sim2real improvement (3DDA) 💥 2–3.5× boosts for π₀ and ACT PEEK gives visuomotor policies the minimal cues to generalize across new tasks, scenes, and semantics.

📄 Paper: 💻 Demo: 🔗 Website, code, data, models: Work done w/ @memmelma (co-lead), @minjunkevink, @fox_dieter17849 , @_jessethomason_ , Fabio Ramos, @ebiyik_ , @abhishekunique7 , @AnqiLi24

Finally, see @memmelma's tweet thread to learn how the method works!
