Which Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encoders
- arXiv: 2608.20280
- Authors: Yash Kulkarni, Shubham Harkare, Arvind Suresh Yogesh Babu
- Published: 2026-08-20; categories: cs.DB, cs.LG (11 pages, 9 figures)
Abstract
Semantic caches reuse an LLM response when the incoming query embedding lies near a cached query, but proposed eviction policies have rarely been compared under one protocol. Using CLEVER, we evaluate FIFO, LRU, LFU, ARC, GDSF, a single-pass streaming adaptation of SISO, and a semantic-redundancy policy across three ordered, deduplicated query corpora, three cache capacities, and two encoders. No evaluated policy improves on LFU by more than 0.041 percentage points in any of the eighteen settings. Replacement is not irrelevant: FIFO and streaming SISO trail LFU by as much as 8.67 and 8.55 points, respectively, at tight capacity. We explain the missing upside with a conditional packing result. Under exact lookup and insert-on-miss, a newly inserted entry cannot have a resident neighbor within the hit radius, so a geometry-aware eviction rule receives little new redundancy signal. A separate audit exposes a larger problem with the evaluated operating point. At MiniLM's median nearest-neighbor threshold, only 2.1-3.9% of sampled LMSYS and QQP hits are judged answer-substitutable, reducing raw hit rates of 51-60% to quality-adjusted rates of 1.1-2.2%. The cross-encoder study further shows that thresholds do not transfer between embedding models. LFU is the strongest simple default in this protocol; deployment decisions should first establish answer validity and then test sub-point policy differences with exact search.
为什么值得读(AAIF 扫描)
用统一评测协议 CLEVER 对语义缓存(semantic cache)的淘汰策略做了一次系统对比:FIFO、LRU、LFU、ARC、GDSF、流式 SISO、语义冗余度策略,跨 3 个有序去重查询语料、3 种缓存容量、2 种编码器,共 18 个设定。结论很干脆:没有任何策略比 LFU 高出 0.041 个百分点以上;但替换策略也不是无关紧要——容量紧张时 FIFO 和流式 SISO 最多落后 LFU 8.67 / 8.55 个百分点。
两个更重要的发现:
1. 为什么花哨策略没赢:精确查找 + 未命中即插入的设定下,新插入的条目在命中半径内不可能有已驻留邻居(条件打包结果),几何感知的淘汰规则根本拿不到新的冗余信号。 2. 命中率的水分:MiniLM 中位最近邻阈值下,LMSYS/QQP 抽样命中的回答可替代率只有 2.1–3.9%,原始命中率 51–60% 缩水成质量调整后的 1.1–2.2%。而且阈值不能跨嵌入模型迁移——换 encoder 就得重标。
工程结论:LFU 是最稳的简单默认值;部署语义缓存应先验证命中答案的有效性,再用精确搜索去测那些不到 1 个百分点的策略差异。做 RAG 缓存选型前必读。
Source: https://arxiv.org/abs/2608.20280
Captured: 2026-08-22 (AAIF content-fetcher)