SGD-KV: Summarization Guided KV Cache Compression
Source: https://arxiv.org/abs/2609.03235
Authors: Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki
Published: 2026-09-03
Categories: cs.CL
PDF: https://arxiv.org/pdf/2609.03235v1
Abstract
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We present SGD-KV (Summarization-Guided KV Cache Compression), a head-aware framework that leverages a novel chunk-summarization diagnostic task to systematically identify and prioritize attention heads specialized in hierarchical information aggregation. Experiments on Qwen2.5-7B-1M and Qwen3-32B across diverse long-context benchmarks demonstrate that SGD-KV achieves state-of-the-art performance with contexts up to 1M tokens, while reducing KV cache memory usage by up to 75%. Our findings show that strategically allocating the KV cache budget based on the summarization score distribution of attention heads yields a superior efficiency-accuracy trade-off for long-context inference.
中文概要
Zeyu Liu 等 2026-09-03 提交,NeurIPS 2026 Efficient Reasoning Workshop 收录。现有 KV cache 压缩方法大多用简单启发式,没有区分不同注意力头的功能差异。SGD-KV 设计了一个 chunk-summarization 诊断任务,系统识别专门做层次信息聚合的注意力头,按头分配预算:算总览的头吃满 cache,普通头压紧。在 Qwen2.5-7B-1M 和 Qwen3-32B 上 1M token 上下文实现 SOTA,KV cache 内存用量最多降到 25%(即压缩 75%)。端侧长上下文要落地,把预算粒度从层细到头是必然的一步。
一句话
SGD-KV 用 chunk-summarization 评分给注意力头分配 KV 预算,1M token 上下文内存最多减 75%,端侧长上下文实用组合的关键一块