正在加载视频...
视频加载失败
LLM inference speed with vs. without KV caching: (learn KV cache management in LLMs below)
8 条评论

Useful taxonomy. The real control is not “KV cache on or off.” It is the model’s cache geometry, the serving engine’s implementation, context and concurrency, cache dtype, paging and eviction policy, offload path, and actual memory topology. We just shipped Aperture Recipe Lab to translate published inference recipes into capability checks for the hardware someone actually owns. Which KV-cache recipe should we encode first?

Left side finished the story and went outside. Right side is still explaining that the little ones were playing. KV cache is just the model finally taking notes.

Good framing for LLM research work. The practical part is not just trying a stronger model, but logging baselines, data splits, task-specific metrics, and failure cases so the result is reproducible.

at the short term cost of everyone iPhone is $100-200 more expensive

Continuous batching changes KV cache overhead significantly, was this measured with static batches or dynamic token-level batching? The difference can be 3-5x on real workloads.

how much of the overhead is from memory access vs actual computation?

Good question. Actually, there is no fixed percentage. It depends on context length, batch size, attention type, precision, GPU, and the serving engine. But during token-by-token decoding, KV-cache attention is usually limited more by memory bandwidth than computation. Prefill is a different phase and is generally more compute-heavy (covered here: At shorter contexts, model-weight reads, projections, MLP layers, communication, and kernel overhead can matter more. As the context grows, KV traffic grows linearly and can dominate. This is why GQA, KV quantization, eviction, and sparse attention improve more than capacity. They reduce the number of bytes each decoding step has to move.

kv caching. you don't respect it til you've run inference without it once. night and day
