When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
原文链接: https://arxiv.org/abs/2609.28870
作者: Yiyu Liu et al.
发布时间: 2026-09-24
源: arxiv
摘要
arXiv 2609.28870 拿两家公司生产 trace 评估了 14 种驱逐算法,在 HBM 限制与大内存池两种设置下都跑。结论直接:相比 Belady oracle 仍有大缺口,但花哨策略相对 LRU 的提升很小。原因在结构:前缀复用被活跃会话的规律节奏主导,recency 异常可预测。论文引入 compute-savings ratio 与两个 offline oracle 来量化这种效应,并指出 prefix cache 还引入两个新挑战:重尾 session footprint 与 attention 计算随序列增长带来高方差 miss cost。直接的反证:对 LLM prefix cache「该不该上复杂策略」是直接答案——大多数情况下 LRU 已经够了。
English Summary
arXiv 2609.28870 evaluates 14 eviction algorithms on production traces from two companies under both HBM-constrained and large memory-pool settings. The conclusion is direct: there is still a large gap to a Belady oracle, but sophisticated policies yield little improvement over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. The paper introduces the compute-savings ratio and two offline oracles to quantify this, and notes two new challenges introduced by prefix caching: heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. The single most leveraged observation is the direct negative result: for LLM prefix cache, 'do we need a fancy policy?' already has its answer — in most settings LRU is enough.
为什么值得关注
主题线扩展 AAIF prefix-caching/eviction/lru/production-traces 等主题。