基础设施 4.0 · 优秀 2026-09-25 · 论文

The KV Cache Is the New Memory Wall

SoK 论文(arXiv 2609.30854,28p)指出长上下文自回归 LLM 推理受内存带宽约束而非算力,绑定资源随序列增长从权重转移到 KV cache以 Llama-3-70B BF16 为例:140 GB 权重超出单卡 80 GB HBM,128k token 单序列再加 42 GB KV论文在 NVIDIA H100 / B200AMD MI300X 三种硬件上推导出随 context 衰减的算术强度闭式解,统一比较量化/驱逐/paging/前缀复用/异构分层五域核心结论是三段式:交叉点以下权重流量主导,KV 压缩几乎无收益;交叉点以上各域都在拿质量换带宽;paging/前缀共享无损但只解容量不解带宽

打开原文回到归档

The KV Cache Is the New Memory Wall

原文链接: https://arxiv.org/abs/2609.30854
作者: Tejinder Singh
发布时间: 2026-09-25
源: arxiv

摘要

SoK 论文(arXiv 2609.30854,28p)指出长上下文自回归 LLM 推理受内存带宽约束而非算力,绑定资源随序列增长从权重转移到 KV cache。以 Llama-3-70B BF16 为例:140 GB 权重超出单卡 80 GB HBM,128k token 单序列再加 42 GB KV。论文在 NVIDIA H100 / B200、AMD MI300X 三种硬件上推导出随 context 衰减的算术强度闭式解,统一比较量化/驱逐/paging/前缀复用/异构分层五域。核心结论是三段式:交叉点以下权重流量主导,KV 压缩几乎无收益;交叉点以上各域都在拿质量换带宽;paging/前缀共享无损但只解容量不解带宽。

English Summary

This SoK paper (arXiv 2609.30854, 28p) frames long-context autoregressive LLM inference as bound by memory bandwidth rather than arithmetic throughput, with the binding resource shifting from weights to the KV cache as sequence length grows. Llama-3-70B in BF16 is the canonical example: 140 GB of weights exceed a single accelerator's 80 GB HBM, and a 128k-token sequence adds 42 GB of KV. The paper derives a closed-form arithmetic intensity as a decaying function of context length, parameterized for NVIDIA H100, NVIDIA B200, and AMD MI300X including per-die bandwidth partitioning and the crossover point. It unifies the five domains of quantization, eviction, paging, prefix reuse, and heterogeneous layering. The single most leveraged observation is a three-part split: below the crossover point, weight traffic dominates and KV compression yields almost nothing; above it, every domain trades quality for bandwidth; and paging/prefix sharing are lossless but only solve capacity, not bandwidth.

为什么值得关注

主题线扩展 AAIF kv-cache/memory-wall/sok/arithmetic-intensity 等主题。

信息源