Rudy Gilman's banner
Rudy Gilman's profile picture

Rudy Gilman

@rgilman333,140 subscribers

substrate agnostic

Shorts

The attention layers in the VAEs for FLUX, Stable Diffusion 3.5, and SDXL don't do anything. You can ablate them with almost no effect. At first I thought they might be involved in some clever circuitry—maybe moving global information—but no they're just flailing around doing nothing.

The attention layers in the VAEs for FLUX, Stable Diffusion 3.5, and SDXL don't do anything. You can ablate them with almost no effect. At first I thought they might be involved in some clever circuitry—maybe moving global information—but no they're just flailing around doing nothing.

88,587 次观看

Half the channels in the last layer of the sdxl vae are disabled intentionally by the model itself. I've never seen this type of dead neuron before—it's like a form of apoptosis where the model kills off a part of itself so the rest may thrive. This is the final conv layer where we produce the output RGB. Notice half the neurons never contribute. The weights on the ignored channels are tiny.

Half the channels in the last layer of the sdxl vae are disabled intentionally by the model itself. I've never seen this type of dead neuron before—it's like a form of apoptosis where the model kills off a part of itself so the rest may thrive. This is the final conv layer where we produce the output RGB. Notice half the neurons never contribute. The weights on the ignored channels are tiny.

59,955 次观看

The sdxl-VAE models a substantial amount of noise. Things we can't even see. It meticulously encodes the noise, uses precious bottleneck capacity to store it, then faithfully reconstructs it in the decoder. I grabbed what I thought was a simple black vector circle on a white background but the VAE latched on to a veritable Navajo quilt of background noise. I was absolutely flummoxed until I realized what was going on! A lot of capacity being used on things we don't need to be modelling at all. Suspected culprit is group norm.

The sdxl-VAE models a substantial amount of noise. Things we can't even see. It meticulously encodes the noise, uses precious bottleneck capacity to store it, then faithfully reconstructs it in the decoder. I grabbed what I thought was a simple black vector circle on a white background but the VAE latched on to a veritable Navajo quilt of background noise. I was absolutely flummoxed until I realized what was going on! A lot of capacity being used on things we don't need to be modelling at all. Suspected culprit is group norm.

52,110 次观看

One of the reasons I like distillation is because you can use it to combine the strengths of different models. RADIO from Pavlo Molchanov and the team at nvidia has the smooth, rich feature maps of DINO-v2 and it can read like Siglip / CLIP. RADIO even uses TimDarcet's registers to consolidate global information and keep the spatial features clean! I think we'll see lots more of this approach in the future. Here's the smoothness you get from distilling DINO-v2 and adding registers. No spatial artifacts.

One of the reasons I like distillation is because you can use it to combine the strengths of different models. RADIO from Pavlo Molchanov and the team at nvidia has the smooth, rich feature maps of DINO-v2 and it can read like Siglip / CLIP. RADIO even uses TimDarcet's registers to consolidate global information and keep the spatial features clean! I think we'll see lots more of this approach in the future. Here's the smoothness you get from distilling DINO-v2 and adding registers. No spatial artifacts.

43,046 次观看

The later features in DINO-v2 are more abstract and semantically meaningful than I'd expected from the training objectives. This neuron responds only to hugs. Nothing else, just hugs.

The later features in DINO-v2 are more abstract and semantically meaningful than I'd expected from the training objectives. This neuron responds only to hugs. Nothing else, just hugs.

34,378 次观看

This is Siglip-2's dedicated DEI neuron. It fires for LGBTQ and indigenous flags, BLM imagery, and especially mixed-race groups doing happy things together (e.g. business meetings, jumping triumphantly, reaching summits etc)

This is Siglip-2's dedicated DEI neuron. It fires for LGBTQ and indigenous flags, BLM imagery, and especially mixed-race groups doing happy things together (e.g. business meetings, jumping triumphantly, reaching summits etc)

23,092 次观看

A challenge for ML diagnosticians: The patient, Segment-Anything-Model-2 (SAM-2), is presenting with two prominent symptoms: Symptom 1) Extremely high-magnitude tokens spread evenly across spatial dimensions.

A challenge for ML diagnosticians: The patient, Segment-Anything-Model-2 (SAM-2), is presenting with two prominent symptoms: Symptom 1) Extremely high-magnitude tokens spread evenly across spatial dimensions.

25,936 次观看

Videos

没有更多内容可加载