Video yükleniyor...
Video Yüklenemedi
🕳️🐇Into the Rabbit Hull – Part I (Part II tomorrow) An interpretability deep dive into DINOv2, one of vision’s most important foundation models. And today is Part I, buckle up, we're exploring some of its most charming features.
64,170 görüntüleme • 11 ay önce •via X (Twitter)
19 Yorum

Assuming the Linear Rep. Hypothesis, SAEs arise naturally as instruments for concept extraction, they will be our companions in this descent. Archetypal SAE uncovered 32k concepts. Our first observation: different tasks recruit distinct regions of this conceptual space.

Let's zoom in on classification. For every class, we find two concepts: one fires on the object (e.g., "rabbit"), and another fires everywhere *except* the object -- but only when it's present! We call them Elsewhere Concepts (credit: @davidbau).

This kind of concept breaks a key assumption in interpretability: that a concept is about the tokens where it fires. Here it is the opposite—the concept is defined by where it does not fire. An open question is how models form such concepts.

Another surprise here: the most important concepts are not object-centric at all, but boundary detectors. Remarkably, these concepts coalesce into a low-dimensional subspace within (see paper).

Now for depth estimation. How does DINO know depth? It turns out it has discovered several human-like monocular depth cues: texture gradients resembling blurring or bokeh, shadow detectors, and projective cues. Most units mix cues, but a few remain remarkably pure.

Curious tokens, the registers. DINO seems to use them to encode global invariants: we find concepts (directions) that fire exclusively (!) on registers. Example of such concepts include motion blur detector and style (game screenshots, drawings, paintings, warped images...)

Huge thanks to all collaborators who made this work possible, and especially to @WangBinxu. This work grew from a year of collaboration! Tomorrow, Part II: geometry of concepts and Minkowski Representation Hypothesis. 🕹️ 📄

@Napoolar Wooow amazing work! is it possible to apply it with a fine tuned model?

@Napoolar @akjags Awesome stuff!

@akjags Thx Nikhil ! Same goes for you 😉

@Napoolar 💪

Thx Kosta 😉 !!

@Napoolar Congrats Thomas! Incredible work!!!

Thx Akshay 🤠

@Napoolar Amazing work, @Napoolar! Looking forward to diving into the paper this weekend!

@Napoolar Looking forward to Part II! It's exciting to see such an in-depth exploration of DINOv2’s capabilities and interpretability—can’t wait to discover what insights you’ll share next.

@Napoolar Looking forward to this series. DINOv2 remains such a fascinating model for how it captures visual structure without supervision. What part of its interpretability are you most curious to unpack first—feature emergence or layer level abstraction?

@Napoolar Fascinating work.

@Napoolar "Into the Rabbit Hull"? My therapist is going to have *so* many questions. Buckled up and praying for a smooth landing, not another existential crisis. 🐰✨

