基础设施 4.0 · 优秀 2026-02-17 · 论文

Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devi...

边缘设备上多 agent LLM 的显存困境:Apple M4 Pro 10.2GB cache 预算下 FP16 8K 上下文只能容纳 3 个 agent,10-agent 工作流需频繁驱逐重载,无持久化时每次驱逐都触发完整 prefill(4K 上下文每 agent 15.7 秒)方案是把每个 agent 的 KV cache 以 4-bit 量化持久化到磁盘并直接重载进 attention 层,免除重复 O(n) prefill收录理由:把 agent 记忆下沉到 KV cache 层的系统设计,对端侧多 agent 部署有直接工程价值

打开原文回到归档

Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices

中文导读

边缘设备上多 agent LLM 的显存困境:设备 RAM 装不下所有 agent 的 KV cache——Apple M4 Pro 在 10.2 GB cache 预算下,FP16 8K 上下文只能容纳 3 个 agent;10-agent 工作流必须频繁驱逐与重载,而无持久化时每次驱逐都触发完整 re-prefill(4K 上下文每 agent 15.7 秒)。方案:把每个 agent 的 KV cache 以 4-bit 量化格式持久化到磁盘,重载时直接注入 attention 层,通过直接 cache 恢复免除冗余的 O(n) prefill 计算。系统三组件:per-agent 隔离 Q4 KV cache 的 block pool(safetensors 格式);跨多个 agent 量化 cache 并发推理的 BatchQuantizedKVCache;跨阶段累积 attention 状态、无需重算的 cross-phase context injection。在三种架构上评估(Gemma 3 12B dense GQA 48 层;DeepSeek-Coder-V2-Lite 16B MoE MLA 27 层;Llama 3.1 8B dense GQA 32 层):cache 恢复使 time-to-first-token 最高降低 136×(Gemma 4K–32K 为 22–136×;DeepSeek 11–76×;Llama 24–111×,1K 时 3–10×);Q4 量化让固定内存可容纳 4× 于 FP16 的 agent 上下文;使用真实 Q4 KV cache 测得的困惑度变化为 Gemma −0.7%、Llama +2.8%、DeepSeek +3.0%。

为什么值得关注

把 agent 记忆下沉到 KV cache 层的系统设计,对端侧多 agent 部署有直接工程价值:不改变 prompt 层的记忆抽象,而是让"换入换出 agent"的成本从整段 prefill 变为一次磁盘读取 + 反量化。三种架构(dense GQA / MoE MLA)上的 TTFT 加速区间与真实 Q4 cache 的困惑度代价数据,是选型时少见的 grounded 数字。开源实现:github.com/yshk-mxim/agent-memory。

关键信息

  • 论文标题:Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices
  • arXiv:2603.04428(v1,2026-02-17);24 页 / 6 图 / 16 表
  • 主分类:cs.LG;跨类:cs.AI
  • 开源实现:https://github.com/yshk-mxim/agent-memory
  • 硬件设定:Apple M4 Pro,10.2 GB KV cache 预算
  • 三组件:block pool(per-agent Q4 cache,safetensors)、BatchQuantizedKVCache、cross-phase context injection
  • 评估架构:Gemma 3 12B(dense GQA, 48L)、DeepSeek-Coder-V2-Lite 16B(MoE MLA, 27L)、Llama 3.1 8B(dense GQA, 32L)
  • 关键数字:TTFT 加速 3–136×(视架构与上下文长度);Q4 容量收益 4×;PPL 变化 −0.7% / +2.8% / +3.0%

English Abstract

Multi-agent LLM systems on edge devices face a memory management problem: device RAM is too small to hold every agent's KV cache simultaneously. On Apple M4 Pro with 10.2 GB of cache budget, only 3 agents fit at 8K context in FP16. A 10-agent workflow must constantly evict and reload caches. Without persistence, every eviction forces a full re-prefill through the model -- 15.7 seconds per agent at 4K context. We address this by persisting each agent's KV cache to disk in 4-bit quantized format and reloading it directly into the attention layer, eliminating redundant O(n) prefill computation via direct cache restoration. The system comprises three components: a block pool providing per-agent isolated Q4 KV caches in safetensors format, a BatchQuantizedKVCache for concurrent inference over multiple agents' quantized caches, and cross-phase context injection that accumulates attention state across conversation phases without re-computation. Evaluated on three architectures (Gemma 3 12B, dense GQA, 48 layers; DeepSeek-Coder-V2-Lite 16B, MoE MLA, 27 layers; Llama 3.1 8B, dense GQA, 32 layers), cache restoration reduces time-to-first-token by up to 136x (Gemma: 22--136x at 4K--32K; DeepSeek: 11--76x at 4K--32K; Llama: 24--111x at 4K--16K; 3--10x at 1K). Q4 quantization fits 4x more agent contexts into fixed device memory than FP16. Perplexity measured with actual Q4 KV caches shows -0.7% for Gemma, +2.8% for Llama, and +3.0% for DeepSeek. Open-source at https://github.com/yshk-mxim/agent-memory

Obsidian Notes

  • 内容由 opencli arxiv paper 2603.04428 -f json 拉取的 arXiv 元数据与摘要生成;页数/图表数来自 arXiv comment 字段。
  • 中文导读与价值判断锚定在摘要声明的设定与数字上;未补充摘要之外的实验细节。
  • 相关条目:ce60e01f(Agent Memory 全景综述——本文是"记忆基质下沉到 KV cache 层"的具体案例)。