正在加载视频...

视频加载失败

🕳️🐇Into the Rabbit Hull – Part I (Part II tomorrow) An interpretability deep dive into DINOv2, one of vision’s most important foundation models. And today is Part I, buckle up, we're exploring some of its most charming features.

64,170 次观看 • 11 个月前 •via X (Twitter)

19 条评论

Thomas Fel 的头像
Thomas Fel11 个月前

Assuming the Linear Rep. Hypothesis, SAEs arise naturally as instruments for concept extraction, they will be our companions in this descent. Archetypal SAE uncovered 32k concepts. Our first observation: different tasks recruit distinct regions of this conceptual space.

Thomas Fel 的头像
Thomas Fel11 个月前

Let's zoom in on classification. For every class, we find two concepts: one fires on the object (e.g., "rabbit"), and another fires everywhere *except* the object -- but only when it's present! We call them Elsewhere Concepts (credit: @davidbau).

Thomas Fel 的头像
Thomas Fel11 个月前

This kind of concept breaks a key assumption in interpretability: that a concept is about the tokens where it fires. Here it is the opposite—the concept is defined by where it does not fire. An open question is how models form such concepts.

Thomas Fel 的头像
Thomas Fel11 个月前

Another surprise here: the most important concepts are not object-centric at all, but boundary detectors. Remarkably, these concepts coalesce into a low-dimensional subspace within (see paper).

Thomas Fel 的头像
Thomas Fel11 个月前

Now for depth estimation. How does DINO know depth? It turns out it has discovered several human-like monocular depth cues: texture gradients resembling blurring or bokeh, shadow detectors, and projective cues. Most units mix cues, but a few remain remarkably pure.

Thomas Fel 的头像
Thomas Fel11 个月前

Curious tokens, the registers. DINO seems to use them to encode global invariants: we find concepts (directions) that fire exclusively (!) on registers. Example of such concepts include motion blur detector and style (game screenshots, drawings, paintings, warped images...)

Thomas Fel 的头像
Thomas Fel11 个月前

Huge thanks to all collaborators who made this work possible, and especially to @WangBinxu. This work grew from a year of collaboration! Tomorrow, Part II: geometry of concepts and Minkowski Representation Hypothesis. 🕹️ 📄

Sebastian A. 的头像
Sebastian A.11 个月前

@Napoolar Wooow amazing work! is it possible to apply it with a fine tuned model?

Nikhil Parthasarathy 的头像
Nikhil Parthasarathy11 个月前

@Napoolar @akjags Awesome stuff!

Thomas Fel 的头像
Thomas Fel11 个月前

@akjags Thx Nikhil ! Same goes for you 😉

Kosta Derpanis at #ECCV2026 🇸🇪 的头像
Kosta Derpanis at #ECCV2026 🇸🇪11 个月前

@Napoolar 💪

Thomas Fel 的头像
Thomas Fel11 个月前

Thx Kosta 😉 !!

Akshay Jagadeesh 的头像
Akshay Jagadeesh11 个月前

@Napoolar Congrats Thomas! Incredible work!!!

Thomas Fel 的头像
Thomas Fel11 个月前

Thx Akshay 🤠

Ali Shehral 的头像
Ali Shehral11 个月前

@Napoolar Amazing work, @Napoolar! Looking forward to diving into the paper this weekend!

Dustin 的头像
Dustin11 个月前

@Napoolar Looking forward to Part II! It's exciting to see such an in-depth exploration of DINOv2’s capabilities and interpretability—can’t wait to discover what insights you’ll share next.

Agam Chaudhary 的头像
Agam Chaudhary11 个月前

@Napoolar Looking forward to this series. DINOv2 remains such a fascinating model for how it captures visual structure without supervision. What part of its interpretability are you most curious to unpack first—feature emergence or layer level abstraction?

Out of service 的头像
Out of service11 个月前

@Napoolar Fascinating work.

StupidFood 的头像
StupidFood11 个月前

@Napoolar "Into the Rabbit Hull"? My therapist is going to have *so* many questions. Buckled up and praying for a smooth landing, not another existential crisis. 🐰✨

相关视频