The KV Cache Working Set: Online Capacity Planning for LLM Inference Systems
原文链接: https://arxiv.org/abs/2609.27746
作者: Luchang Li et al.
发布时间: 2026-09-23
源: arxiv
摘要
KVSET(arXiv 2609.27746)提出一个在线分析器,用 Mattson 栈算法估计 LLM serving 负载的「KV cache 工作集」——即达到目标命中率所需的最小容量——避免逐容量模拟。前缀缓存在 agentic 工作负载上特别关键,因为模型会被反复调用并持续追加会话/工具历史;保留全部历史 KV 太贵,容量不够又显著降低命中率。KVSET 与 ActKV 互补——一个管驱逐策略,一个管容量规划。
English Summary
KVSET (arXiv 2609.27746) is an online analyzer that estimates the KV cache 'working set' of LLM serving workloads — the minimum cache capacity required to hit a target hit rate — using the Mattson stack algorithm, avoiding per-capacity simulation. Prefix caching matters most for agentic workloads, where models are repeatedly invoked with growing conversation and tool-use history: keeping all historical KV states is prohibitively expensive, but insufficient capacity degrades hit rates substantially. KVSET complements ActKV: one decides the eviction policy, the other decides how much capacity to provision.
为什么值得关注
主题线扩展 AAIF kv-cache/prefix-caching/capacity-planning/mattson-stack 等主题。