Video yükleniyor...
Video Yüklenemedi
Diffusion models spend the same compute on a blank wall as on a face. But you often know in advance where the detail will be. Introducing Level-of-Token (LoT) Diffusion: we turn that knowledge into a multiresolution token layout, with fine tokens where detail is needed and coarse tokens elsewhere. 1/9🧵
40,175 görüntüleme • 1 gün önce •via X (Twitter)
19 Yorum

Our key idea: generalize the uniform token grid of pretrained DiTs to a Level-of-Token layout. Each token is a rectangle of any size, and together they tile the whole image. 2/9

How it works: Level-of-Token DiT minimally modifies a pretrained DiT. Each token is projected from the latent patches it covers, the DiT is conditioned on token shapes, and extent-dependent heads restore the asymmetric velocity, which is converted into the full-rank velocity. We fine-tune both image (Flux.2) and video (Wan2.1) models with patch-wise asymmetric flow matching, so the pretrained generative prior is preserved. 3/9

LoT layouts can come from any signal that says where detail matters. Left to right: bounding boxes, semantic masks, texture variance, and depth, with the source in the corner; each tile starts as its LoT layout, then reveals the generation. One model, any layout source. 4/9

The same holds for video: per-frame LoT layouts follow what moves. Each video starts as its LoT layout that's derived from semantic masks (top left), bounding boxes (top right), texture variance (bottom left), and depth (bottom right), with 1.3–2.1× fewer tokens. 5/9

An agent can also author the layout directly. Here the agent paints an importance map on a canvas (left), marking where detail matters for “a girl riding a corgi”. LoT turns it into a token layout with 2.6× fewer tokens and generates the image (right). 6/9

Going further, an agent can plan detail over time. Here it builds a rough 3D scene of a toy train in Blender. LoT turns the plan into a per-frame token layout (overlaid) and generates the video with 1.5× fewer tokens. 7/9

Generation adapts to the layout: fine tokens go where the prompt needs detail, and the speedup is set by the layout's token budget. Here LoT (right) uses 6,632 tokens instead of 14,336 and generates 2.5× faster than full resolution (left), and up to 4.6× when detail is more concentrated. 8/9

Work led by @GeorgeNaka40190, together with @BrianCChao, @jan_on_x, @HanshengCh, @fedassa, Leonidas Guibas, @YarivLior. 📄 Paper: 🌐 Project page:

Excited to share LoT Diffusion with the amazing team! Beyond efficiency, what I love most is the new kind of control it gives you: where the detail goes. Even a rough painted map or a typography mask can steer the generation. More in my thread 👇

No citation? :( But honestly, really cool work!!! :) Love to see adaptive methods thrive!

Great work! Can’t wait to try it🔥

feels like mipmaps for diffusion, cheap faces finally

Adaptive token allocation could make diffusion models significantly more compute-efficient

coarse tokens everywhere else is the whole trick

algos like this are desperately needed

该细的地方细,算力花对了

2x faster on images and 3.5x on video just by refusing to spend full tokens on blank walls.

this is super exciting! fixing the measurement op. foveated

That's cool! Did this enable higher resolution image/video generations given same compute budget without LoT?
