
Rudy Gilman
@rgilman33 • 3,140 subscribers
substrate agnostic
Shorts
Videos

DINO-v3 has a single high-magnitude channel on its residual pathway, channel 416. Turning off this single channel affects DINO's entire output by 50-80%. For context, turning off a random channel has an effect of less than one percent. The model builds up channel 416 in its last two layer-scale operations, using a single high-magnitude weight in each op to drastically ramp up channel 416's magnitude. This channel doesn't depend on the input, every image fires with a constant overlay. After bringing channel 416 up to a value of about ten thousand, DINO-v3 then scales it down in the final layer-norm to almost nothing, removing it without a trace.
Rudy Gilman85,984 次观看 • 10 个月前

The VAE used in SDXL has extremely high-magnitude "splotches" in its latents. The individual neurons in these blobs fire with magnitudes of close to a million. These aren't some accident of training or initialization—the model creates these high-magnitude splotches for a specific reason: to circumvent the group-norm operations.
Rudy Gilman104,849 次观看 • 1 年前

This layer in DINO-v2 dedicates about half its attention mass to a single operation. Each of the sixteen heads independently learns the same circuit to perform this task. What is this all-important operation? The “no-op”. That’s right, we’re spending half our computation to do… absolutely nothing.
Rudy Gilman97,147 次观看 • 1 年前

The majority of features in this layer of Siglip-2 are multimodal. I'd expected some multimodality but was surprised that two-thirds of the neurons I tested bind together their visual and linguistic features. This neuron fires for images of mustaches and for the word "mustache"
Rudy Gilman18,519 次观看 • 1 年前

Through the length of SDXL-Turbo you can see the image forming in the activations like an object taking shape in the afternoon clouds. Like David emerging from a block of granite. At two-thirds through the model you can see the image clearly. It's spatially crisp, refined. Then in the last third of the model something strange but predictable happens: the model translates this well-formed image back into noise. By the output of the U-net, the activations are again completely uninterpretable. Why does this happen? Because we're asking the model to predict the noise rather than the clean image. This is awkward. Like Michelangelo visualizing the chips of granite surrounding the David rather than the David itself. Like asking a friend to look at a cloud and imagine not a shape, but the extra bits of cloud that when removed will leave a shape remaining.
Rudy Gilman11,641 次观看 • 1 年前
没有更多内容可加载