Загрузка видео...

Не удалось загрузить видео

На главную

Introducing a new KV format: VBR (Variable Bit Rate) VBR dynamically quantizes your KV cache layer-by-layer as your session grows, giving you the highest quality possible within your VRAM constraints. This is my dream format. Pushed to master. Available now. 1/15 🧵

48,983 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 14

Фото профиля buun
buun3 месяцев назад

Mixed codec caches have thus far been very crude. Either assymetric cache like -ctk turbo8 -ctv turbo4 (6.125 BPV), or Sparse-V which naively protects the top couple of layers. We have profiled every layer of Qwen & Gemma for what the actual, measured best mix order is.

Фото профиля buun
buun3 месяцев назад

But the beauty of VBR is that, say the only way to fit 150k context is with turbo4. With VBR, you don't *actually* have a cache that quantized until you've literally hit 150k tokens. Your agent will be F16 until 18k, then F16/T8 mixes, and so on.

Фото профиля buun
buun3 месяцев назад

If you never want to run something as quantized as 1-bit, you can set the --vbr-floor to the lowest quality you can tolerate. You can also set the VRAM you want to allocate to your kv with --vbr-vram. VBR will give you the best possible quality for whatever your constraints are.

Фото профиля buun
buun3 месяцев назад

Major Codec Upgrades- buun-llama has significantly upgraded its underlying codecs with this release. Thanks to research by @tbraun96 (TorQuant), we have shaved an additional 40% off KLD of all turbo types (!), and, new advancements in our TCQ have given us an additional -20%.

Фото профиля buun
buun3 месяцев назад

All advancements in the underlying KV codecs are rolled into VBR. You no longer have to learn the new flags for "TorQuant", "KVaRN", "UberQuant Pro+". We will keep researching and developing, and -vbr will always give you the best possible quality.

Фото профиля buun
buun3 месяцев назад

More fun new things in this release: MTP support for Gemma4 12B, DFlash support for Gemma4 31B, fused kernels have been ported and optimized to ROCm - AMD users are no longer second class citizens on our fork. We have rebased to upstream as of 7/8/26. Improvements in PP speed across the board. And much more.

Фото профиля Gabriel A. Devenyi
Gabriel A. Devenyi3 месяцев назад

How will this interact with `--fit` in the case of VRAM constrained MoE inference?

Фото профиля buun
buun3 месяцев назад

It's a great combo with --fit and MoEs. It will overflow the sparse expert weights to CPU, keep attention layers on GPU, and use the remaining vram for the dynamic VBR. VBR works perfectly with GPU sharding etc aswell. One note is that if any KV has to offload to CPU, it will offload in plain q8, while any layers remaining on GPU will be VBR.

Фото профиля Bruno Vilela
Bruno Vilela2 месяцев назад

With VBR dynamically quantizing the KV cache layer-by-layer as context grows, would it also be feasible to reserve a portion of the KV space (or manage extra slots) for injecting and removing custom K/V vectors in real time during generation? The goal would be to use part of the cache as a kind of hidden scratchpad that is sturdy enought to survive quantization.

Фото профиля buun
buun2 месяцев назад

I think that this is possible without VBR, and could just work along side it (injected KV would push a degrade, just like any additional KV does). I'm going to be doing some work in this direction for sure, have you seen LatentMAS?

Фото профиля DesignCntrl Inc. / Destrozado
DesignCntrl Inc. / Destrozado2 месяцев назад

KV direct

Фото профиля Nick
Nick2 месяцев назад

this is actually huge for long context sessions, been waiting for something like this

Фото профиля Taeyun Jang
Taeyun Jang2 месяцев назад

좋은 아이디어 공유 감사합니다.

Фото профиля Felipe Sztutman
Felipe Sztutman2 месяцев назад

ran your VBR/TCQ through my decoy-at-depth probe: exact recall at depth, not perplexity. turbo2_tcq recovers 2-5x better than plain 2-bit at 32k, same two bits and a better trellis codebook. and the keys-first split as a fingerprint: 2-bit V is nearly free, 2-bit K is the cost.

Похожие видео

I just crammed the updated Gemma 4 26B A4B QAT (MoE) with 180k context into an 8GB RTX 4060 (8 GB VRAM + 16 GB RAM only!!) and optimized the batch size. 23 tokens/sec decode, 300 tokens/sec prefill Yesterday I showed you a Gemma 4 31B dense model running flawlessly on an RTX 4090. Today, we're breaking the VRAM bank on a budget card using Unsloth’s new Gemma 4 26B (A4B) QAT quants. Following Google’s chat template update that boosted agentic benchmarks by +10%, I pushed this model to its absolute limits. Here is how you squeeze 250k context out of 8GB of VRAM. # The Setup & The Optimization - Hardware: Nvidia RTX 4060 (8GB VRAM) + 16GB System RAM - Environment: CUDA 13.0 build of llama.cpp - Model: gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf - Prompt: 28,000 tokens of prompt for each run If you read my L2 cache breakdown (attached in replies), you know the 4060’s 24MB cache maxes out at `-b 1024 -ub 1024`. Push past that, and prefill crashes. I locked those flags in for every test below to ensure maximum GEMM throughput. # 1. The Raw Context Push (Unquantized KV Cache) First, I wanted to see how far pure 8GB VRAM + 16GB RAM could stretch without touching the KV cache: - 80k Context: Prefill 385 t/s | Decode 25.5 t/s - 120k Context: Prefill 270 t/s | Decode 24 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 120000 --port 8080 -ub 1024 -b 1024 Without KV quantization, 120k is your hard ceiling. push past that prefill throughput drops off a cliff, making the model practically unusable for large agentic workloads. # 2. The Q8 KV Cache Lifeline To survive 250k context on a budget card, you have to quantize the KV cache. I enabled 8 bit KV cache (`-ctk q8_0 -ctv q8_0`) and re ran: - 180k Context: Prefill 280 t/s | Decode 22.8 t/s - 250k Context: Prefill 115 t/s | Decode 20 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 180000 --port 8080 -b 1024 -ub 1024 -ctk q8_0 -ctv q8_0 Result: Q8 KV cache brings 250k context back from the dead. Decode speed stabilizes at a highly usable 20 t/s. You are trading a very small bit amount of reasoning precision for an extra 130,000 tokens of context window. if you own a single rtx 3050, 3060, 3070, 4050, 4060, 5050 or 5060, you must try this model and optimize your batch size for higher prefill. Hugging Face links to the updated Unsloth's QAT quants and performance graph are in the replies below. What model are you running on your 6GB, 8GB or 12GB cards right now? Let's see your setups.

Alok

36,617 просмотров • 2 месяцев назад