Video yükleniyor...
Video Yüklenemedi
Recent KV cache compression research has focused on small models with dense attention. We scaled these techniques to GLM 5.3, a frontier sparse attention model. We adapted KVzip and Attention Matching, reducing KV cache size by 80% while retaining over 90% of full-context accuracy on the QuALITY benchmark.
18,884 görüntüleme • 10 gün önce •via X (Twitter)
9 Yorum

The original KVzip and Attention Matching implementations use dense attention to score all previous tokens and select a separate subset of keys for each layer-head pair. GLM’s architecture prevents both. Its shared latent cache requires a common selection across heads, while sparse attention scores only the positions its indexer chooses to read.

For KVzip, we use dense scoring and compare shared selection rules that balance importance across heads. For Attention Matching, we compare HAK and OMP selection, then fit per-head biases and shared latent values against GLM’s sparse attention. We fit one layer at a time, accounting for changes introduced by earlier compression.

Read the full writeup for how we adapted these methods to GLM 5.3, what we learned from comparing selection rules, and how accuracy changes as we compress more of the cache.

I think we all know what this means

Scaling these techniques is important to allow them to work in actual production workflows, thank you Jeff Zheng

Time to revert the hype from Jev to Jeff

Did you see any noticeable decode throughput penalties from computing that shared selection across heads?

疎な注意機構のGLM 5.3でもKVキャッシュを80%削れたのは驚きました。小さな密なモデルでの研究が、ここまで届くんですね。

80% smaller KV with 90% QuALITY is wild on sparse GLM. my bet is QuALITY MC over a frozen doc understates the hit once agent tools keep rewriting state and those survivors were scored for a static read
