Loading video...
Video Failed to Load
(1/5) FP4 hardware is here, but 4-bit attention still kills model quality, blocking true end-to-end FP4 serving. To fix that, we propose Attn-QAT, the first systematic study of quantization-aware training for attention. The result: FP4 attention quality is comparable to BF16 attention with 1.1x–1.5x higher throughput than SageAttention3 on... show more
38,252 views • 6 months ago •via X (Twitter)
12 Comments

(2/5)Naive QAT breaks when applied to FlashAttention kernels. We found two fixes are needed: 1. Store a small high-precision auxiliary output so the gradient computation stays mathematically consistent 2. Recompute attention probabilities in the backward pass using the same low precision as the forward pass These two changes stabilize 4-bit attention training.

(3/5) Across both video diffusion models and language models, Attn-QAT recovers the quality drop of 4-bit attention without the extra outlier-mitigation heuristics. For continued pretraining, Attn-QAT recovers most of the quality loss caused by FP4 attention on Qwen3-14B and partially recovers it on Llama 3.1-70B. For supervised fine-tuning, Attn-QAT can be used as a drop-in replacement for BF16 attention. On Qwen3-14B, it achieves nearly identical downstream benchmark performance to BF16 attention. On Llama 3.1-70B, it remains close with a small gap. For randomly-selected example videos (generated by Wan-2.1-14B), we see that with Attn-QAT, FP4 attention produces videos comparable to BF16 attention, whereas SageAttention3 produces videos with artifacts.

(4/5) Because the model learns to account for quantization error during QAT, inference needs no extra heuristics! No Q/K smoothing, no two-level quantization. Simpler kernel → faster inference compared to SageAttention3 on an RTX 5090.

(5/5) On a B200, Blackwell's tensor cores are so fast that the softmax now becomes a bottleneck in addition to the GEMMs. Quantizing PV adds scale-factor overhead that piles onto the softmax warps. So we run NVFP4 QK + BF16 PV, with a careful TMEM overlap schedule to fit scale factors into an already-full pipeline. Result: 1801 TFLOPS and up to 1.39x over FlashAttention-4. 2x/4x faster exp on B300/Rubin should push this further, and end-to-end FP4 serving, once blocked by attention quality, is now within reach.

dang i see this on PRs a lot :D tricky now theres all these agents to control what they are all doing, maybe they should set a githook or smth @grok whats best way to prevent this

@gork is this true man

@grok hi

nice work

Great work!

@ye_combinator Did you try FP8? It should be the same performance uplift as FP4

WOW!

friends from @fal might like this :)
