基础设施 4.0 · 优秀 2026-09-25 · 论文

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV(arXiv 2609.31395)是第一个面向 agentic 推理的 KV cache 压缩框架,关键思路是把驱逐标准从整体输出质量改为对动作生成的贡献三点核心:Action-oriented KV cache 驱逐利用稳定的动作访问模式保留对未来动作关键的条目;Confidence-driven adaptive budget allocation 用 LLM 内在置信度把预算分配适配到演化的动作关键内存需求;Page-aware compression 把压缩标准化成三个原语并配 kernel长 trace 任务上,ActKV 用 FullKV 25.98% 的峰值内存保留 98.53% 的精度,token/task 吞吐分别达到 FullKV 的 3.97x / 3.58x

打开原文回到归档

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

原文链接: https://arxiv.org/abs/2609.31395
作者: Zihan Wang et al.
发布时间: 2026-09-25
源: arxiv

摘要

ActKV(arXiv 2609.31395)是第一个面向 agentic 推理的 KV cache 压缩框架,关键思路是把驱逐标准从「整体输出质量」改为「对动作生成的贡献」。三点核心:Action-oriented KV cache 驱逐利用稳定的动作访问模式保留对未来动作关键的条目;Confidence-driven adaptive budget allocation 用 LLM 内在置信度把预算分配适配到演化的动作关键内存需求;Page-aware compression 把压缩标准化成三个原语并配 kernel。长 trace 任务上,ActKV 用 FullKV 25.98% 的峰值内存保留 98.53% 的精度,token/task 吞吐分别达到 FullKV 的 3.97x / 3.58x。

English Summary

ActKV (arXiv 2609.31395) is the first KV cache compression framework tailored for agentic LLM inference. The key idea is to value KV entries by their contribution to action generation rather than overall output quality. Three pillars: action-oriented KV cache eviction that exploits stable action access patterns to retain entries critical to future actions; confidence-driven adaptive budget allocation that uses the LLM's intrinsic confidence to fit the budget to evolving action-critical memory demands; and page-aware compression that standardizes compression into three primitives with custom kernels. On long-trace tasks ActKV retains 98.53% of FullKV accuracy at only 25.98% of its peak KV memory, and reaches 3.97x/3.58x FullKV's token/task throughput.

为什么值得关注

主题线扩展 AAIF kv-cache/agent-inference/eviction/paged-memory 等主题。

信息源