Video yükleniyor...
Video Yüklenemedi
Introducing a new KV format: VBR (Variable Bit Rate) VBR dynamically quantizes your KV cache layer-by-layer as your session grows, giving you the highest quality possible within your VRAM constraints. This is my dream format. Pushed to master. Available now. 1/15 🧵
48,983 görüntüleme • 3 ay önce •via X (Twitter)
14 Yorum

Mixed codec caches have thus far been very crude. Either assymetric cache like -ctk turbo8 -ctv turbo4 (6.125 BPV), or Sparse-V which naively protects the top couple of layers. We have profiled every layer of Qwen & Gemma for what the actual, measured best mix order is.

But the beauty of VBR is that, say the only way to fit 150k context is with turbo4. With VBR, you don't *actually* have a cache that quantized until you've literally hit 150k tokens. Your agent will be F16 until 18k, then F16/T8 mixes, and so on.

If you never want to run something as quantized as 1-bit, you can set the --vbr-floor to the lowest quality you can tolerate. You can also set the VRAM you want to allocate to your kv with --vbr-vram. VBR will give you the best possible quality for whatever your constraints are.

Major Codec Upgrades- buun-llama has significantly upgraded its underlying codecs with this release. Thanks to research by @tbraun96 (TorQuant), we have shaved an additional 40% off KLD of all turbo types (!), and, new advancements in our TCQ have given us an additional -20%.

All advancements in the underlying KV codecs are rolled into VBR. You no longer have to learn the new flags for "TorQuant", "KVaRN", "UberQuant Pro+". We will keep researching and developing, and -vbr will always give you the best possible quality.

More fun new things in this release: MTP support for Gemma4 12B, DFlash support for Gemma4 31B, fused kernels have been ported and optimized to ROCm - AMD users are no longer second class citizens on our fork. We have rebased to upstream as of 7/8/26. Improvements in PP speed across the board. And much more.

How will this interact with `--fit` in the case of VRAM constrained MoE inference?

It's a great combo with --fit and MoEs. It will overflow the sparse expert weights to CPU, keep attention layers on GPU, and use the remaining vram for the dynamic VBR. VBR works perfectly with GPU sharding etc aswell. One note is that if any KV has to offload to CPU, it will offload in plain q8, while any layers remaining on GPU will be VBR.

With VBR dynamically quantizing the KV cache layer-by-layer as context grows, would it also be feasible to reserve a portion of the KV space (or manage extra slots) for injecting and removing custom K/V vectors in real time during generation? The goal would be to use part of the cache as a kind of hidden scratchpad that is sturdy enought to survive quantization.

I think that this is possible without VBR, and could just work along side it (injected KV would push a degrade, just like any additional KV does). I'm going to be doing some work in this direction for sure, have you seen LatentMAS?

KV direct

this is actually huge for long context sessions, been waiting for something like this

좋은 아이디어 공유 감사합니다.

ran your VBR/TCQ through my decoy-at-depth probe: exact recall at depth, not perplexity. turbo2_tcq recovers 2-5x better than plain 2-bit at 32k, same two bits and a better trellis codebook. and the keys-first split as a fingerprint: 2-bit V is nearly free, 2-bit K is the cost.
