基础设施 4.0 · 优秀 2026-09-27 · 论文

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

同一套 coding-agent 负载,KV 卸载到主存在的一个部署加速另一个变慢第三个无感原因是缓存必须活到再次被用:一个 agent 等工具时,服务端在处理其余所有 agent 的上下文,前缀能复用的条件是 host 层装得下整个 agent 池的可复用上下文,作者称为 reuse working set方案用 stack-distance 模型从 agent 历史在线估计 working set 来定 host 层容量;容量不足时停止写入会被驱逐的大块 refill只扩展仍被缓存的前缀,充足时全量写入SWE-bench Verified 上按估计 working set 定容量后重算 prompt token 减少 93%端到端时间减少 39%...

打开原文回到归档

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

  • ID: 49b8a0a8
  • 原文链接: https://arxiv.org/abs/2609.33762
  • 作者: Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, Haoli Bai
  • 日期: 2026-09-27
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, offloading, agent-serving, capacity-planning
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

同一套 coding-agent 负载,KV 卸载到主存在的一个部署加速、另一个变慢、第三个无感——原因是缓存必须活到再次被用:一个 agent 等工具时,服务端在处理其余所有 agent 的上下文,前缀能复用的条件是 host 层装得下整个 agent 池的可复用上下文,作者称为 reuse working set。方案用 stack-distance 模型从 agent 历史在线估计 working set 来定 host 层容量;容量不足时停止写入会被驱逐的大块 refill、只扩展仍被缓存的前缀,充足时全量写入。SWE-bench Verified 上按估计 working set 定容量后重算 prompt token 减少 93%、端到端时间减少 39%;小固定容量下策略仍砍 35% 重算,大容量下避免 always-filter 写入造成的 4.3 倍重算膨胀。卸载划算条件:GPU 每 Byte 主机带宽分到的算力偏低,且 host 装得下 working set。代码开源。

论文信息

  • arXiv ID: 2609.33762
  • 提交日期: 2026-09-27
  • arXiv 分类: cs.DC, cs.LG
  • 作者: Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, Haoli Bai
  • 链接: https://arxiv.org/abs/2609.33762

Abstract(arXiv 原文)

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it speeds up one deployment, slows down another, and changes nothing on a third, even where loading a token back is several times cheaper than recomputing it. The reason is that cached state must survive until it is used again. While one agent waits for its tool, the server processes the contexts of all other agents, so an agent's prefix is reused only if the host tier holds the reusable context of the whole agent pool, which we call the reuse working set. A smaller tier keeps writing state that is evicted before anyone reads it. We present EfficientAgent, which sizes and manages the host tier by this working set. A stack-distance model estimates the working set from agent histories to size the host tier; its predictions, made before the experiments, located the capacity at which offloading starts to pay. When the tier is too small, a runtime policy stops writing large refills of evicted context and keeps extending prefixes that are still cached; when the tier is large enough, it writes everything. On SWE-bench Verified coding agents, a host tier sized to the estimated working set cuts recomputed prompt tokens by 93% and end-to-end time by 39%. With a small fixed tier, the policy cuts recomputation by 35%; with a large tier, it avoids the 4.3-fold increase caused by always filtering writes. Across three GPU types and two models, offloading pays off when the GPU has little compute per byte of host bandwidth and the host tier holds the working set. Code is available at https://github.com/KunmingSHAO/efficientagent_release.

为什么值得关注

卸载划算的条件被写成两个可计算量,working set 估算方法可复用到自己的 serving 容量规划——先算再买内存。

English Summary

Why KV offloading helps one agent deployment, hurts another, and leaves a third unchanged: cached state must survive until reused, so a prefix is only reused if the host tier holds the reusable context of the whole agent pool — the reuse working set. EfficientAgent sizes the host tier with a stack-distance model estimated online from agent histories and filters writes when the tier is too small. On SWE-bench Verified, a working-set-sized tier cuts recomputed prompt tokens 93% and end-to-end time 39%; the write policy saves 35% recomputation with small tiers and avoids a 4.3x blowup with large ones. Offloading pays when GPU compute per byte of host bandwidth is low and the tier holds the working set. Code open source.