Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

LLM inference speed with vs. without KV caching: (learn KV cache management in LLMs below)

41,032 görüntüleme • 22 gün önce •via X (Twitter)

8 Yorum

Jonathan Sandhu profil fotoğrafı
Jonathan Sandhu22 gün önce

Useful taxonomy. The real control is not “KV cache on or off.” It is the model’s cache geometry, the serving engine’s implementation, context and concurrency, cache dtype, paging and eviction policy, offload path, and actual memory topology. We just shipped Aperture Recipe Lab to translate published inference recipes into capability checks for the hardware someone actually owns. Which KV-cache recipe should we encode first?

Misha_Cripto profil fotoğrafı
Misha_Cripto22 gün önce

Left side finished the story and went outside. Right side is still explaining that the little ones were playing. KV cache is just the model finally taking notes.

cordivai | Machine Learning & AI profil fotoğrafı
cordivai | Machine Learning & AI22 gün önce

Good framing for LLM research work. The practical part is not just trying a stronger model, but logging baselines, data splits, task-specific metrics, and failure cases so the result is reproducible.

Sean Dong profil fotoğrafı
Sean Dong20 gün önce

at the short term cost of everyone iPhone is $100-200 more expensive

Tang Vu profil fotoğrafı
Tang Vu22 gün önce

Continuous batching changes KV cache overhead significantly, was this measured with static batches or dynamic token-level batching? The difference can be 3-5x on real workloads.

deezzex profil fotoğrafı
deezzex22 gün önce

how much of the overhead is from memory access vs actual computation?

Avi Chawla profil fotoğrafı
Avi Chawla22 gün önce

Good question. Actually, there is no fixed percentage. It depends on context length, batch size, attention type, precision, GPU, and the serving engine. But during token-by-token decoding, KV-cache attention is usually limited more by memory bandwidth than computation. Prefill is a different phase and is generally more compute-heavy (covered here: At shorter contexts, model-weight reads, projections, MLP layers, communication, and kernel overhead can matter more. As the context grows, KV traffic grows linearly and can dominate. This is why GQA, KV quantization, eviction, and sparse attention improve more than capacity. They reduce the number of bytes each decoding step has to move.

maguyva profil fotoğrafı
maguyva21 gün önce

kv caching. you don't respect it til you've run inference without it once. night and day

Benzer Videolar

Researchers made LLM inference 14x faster and 90% cheaper. The video below depicts the speed up in action. Providers discount cached input tokens by as much as 90% because a cache hit skips prefill compute entirely. For stable system prompts and tool definitions, hit rates of 60 to 85% are achievable, which makes it the highest-leverage inference optimization. But the cost saving only works when the cached text is an exact, byte-for-byte prefix of the new request. If you change one character anywhere before it, the entire cached region is missed. Three common request patterns produce full cache misses: - A query that needs documents A and B together can't reuse B's standalone cache, because those KV entries were computed without A in front of them. - The same three documents retrieved in a different order produce a full cache miss, even though nothing about the documents changed. - In multi-turn conversations, every new turn invalidates whatever was cached beyond the stable prefix. Alibaba's production data did a study on this and found that just 10% of cached KV blocks serve 77% of all cache hits. So most of what gets cached sits in storage and is never used a single time. And the root cause is that KV entries are position-dependent. Each token's KV encodes attention to everything before it, so a cached block is only valid in the exact context it was computed in. There's a second, less discussed problem as well. Cache management runs inside the inference engine's process. Moving KV tensors between GPU, CPU, and disk competes with inference for the same resources. This is why Google's TurboQuant compresses KV caches to 3 bits with no accuracy loss and still causes a 20%+ slowdown when it runs in-process. Fixing both problems means restructuring where caching lives. Cache management moves into its own process, the engine only exchanges block IDs over shared GPU memory, and heavy data movement runs across GPU, CPU, disk, and remote storage in parallel. Non-prefix reuse gets handled by selectively recomputing only the small set of tokens that attend across document boundaries. LMCache is the open-source project (10k+ stars) that implements this exact architecture, and it plugs into vLLM, SGLang, and TensorRT-LLM. The selective recomputation part is implemented in its CacheBlend technique, which makes cached docs in any order and combination, with 2-4x faster multi-document processing. On H200s running Qwen3-235B with 50 concurrent users, LMCache's multiprocess mode delivers 14x faster time-to-first-token and 4x faster decoding compared to in-process caching. GitHub repo: (don't forget to star 🌟) My co-founder wrote a full breakdown of KV cache management. It covers the disaggregated architecture behind the 14x speed up, how CacheBlend preserves generation quality while skipping recomputation, and how to turn every document in a knowledge base into a reusable cached asset. Read it below.

Avi Chawla

30,717 görüntüleme • 2 ay önce

i spent 3 hours finding the sweet spot for hermes 4.3 36B on a single RTX 3090. saving you the trouble, anon. the model is 21.8GB at Q4_K_M. that leaves 2.2GB free on 24GB VRAM. not much room for KV cache. here's what actually happened: 4K: 35.3 tok/s 8K: 35.2 tok/s 16K: 34.8 tok/s 32K: 34.6 tok/s 64K: 6.4 tok/s 128K: 1.9 tok/s flat from 4K to 32K. then it falls off a cliff at 64K. the trick is quantized KV cache. without it you OOM at 16K. with quantized KV cache you get 32K at full speed. all 64 layers on GPU. at 64K something weird happens. ngl 99 (all layers on GPU) = 3.96 tok/s. the KV cache silently spills to CPU. drop to ngl 55 and speed jumps to 6.37. drop to ngl 48 and it gets worse again (3.46). there's an offload sweet spot where you free just enough VRAM for the cache without losing too much compute to PCIe transfers. 128K works at ngl 32 but you're at 1.95 tok/s. half the model on CPU. usable for batch work, not for interactive. the sweet spot command: llama-server -m hermes-4.3-36b-Q4_K_M.gguf -ngl 99 -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 32K context. 34.6 tok/s. all on GPU. this is where dense 36B lives on 24GB. for comparison, qwen 3.5 (35B MoE, 3B active) holds 112 tok/s from 4K all the way to 262K on the same GPU. no speed drop. same total params, completely different architecture. hybrid linear attention means flat context scaling. dense pays for every token in the KV cache. code and quality comparison coming next. fast vs slow generation side by side in the videos below.

Sudo su

52,268 görüntüleme • 7 ay önce