基础设施 3.0 · 值得看 2026-09-29 · 论文

Efficient Agentic LLM Serving over SSD-based Sparse KV Storage

agentic 会话在推理与工具调用间切换历史越积越长;稀疏注意力加 SSD 存 KV 是便宜组合,但稀疏选择依赖推理中的临时中间值,SSD 读被压进关键路径,还要吃碎片访问与读写干扰Janus 聚焦占历史 KV 加载大头的 append prefill:用较早中间值跑模型自带的 KV 选择模块预测需求(无需额外训练),预测读与计算重叠,miss 在 attention 执行前补读保证输出不变;SSD 侧合并相邻读CPU 上把碎片 KV 页打包顺序写读活跃时限制后台写三个模型三条 agentic trace 上 TTFT 较已有最好工作最高快 1.57-3.69 倍(平均 1.22-1.85 倍),decode 效率不掉

打开原文回到归档

Efficient Agentic LLM Serving over SSD-based Sparse KV Storage

  • ID: 205b1a14
  • 原文链接: https://arxiv.org/abs/2609.36938
  • 作者: Wenhao He, Ping Zhang, Xiaohe Hu, Chutian Wang, Jinlong Hou, Yuan Cheng, Peng Sun, Fangcheng Fu
  • 日期: 2026-09-29
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, ssd, sparse-attention, agent-serving
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

agentic 会话在推理与工具调用间切换、历史越积越长;稀疏注意力加 SSD 存 KV 是便宜组合,但稀疏选择依赖推理中的临时中间值,SSD 读被压进关键路径,还要吃碎片访问与读写干扰。Janus 聚焦占历史 KV 加载大头的 append prefill:用较早中间值跑模型自带的 KV 选择模块预测需求(无需额外训练),预测读与计算重叠,miss 在 attention 执行前补读保证输出不变;SSD 侧合并相邻读、CPU 上把碎片 KV 页打包顺序写、读活跃时限制后台写。三个模型、三条 agentic trace 上 TTFT 较已有最好工作最高快 1.57-3.69 倍(平均 1.22-1.85 倍),decode 效率不掉。

论文信息

  • arXiv ID: 2609.36938
  • 提交日期: 2026-09-29
  • arXiv 分类: cs.DC
  • 作者: Wenhao He, Ping Zhang, Xiaohe Hu, Chutian Wang, Jinlong Hou, Yuan Cheng, Peng Sun, Fangcheng Fu
  • 链接: https://arxiv.org/abs/2609.36938

Abstract(arXiv 原文)

Agentic sessions driven by Large language models (LLMs) often alternate between model inference and tool use, accumulating long histories across successive rounds. Serving these sessions efficiently requires reducing attention computation and retaining history key-value (KV) caches to avoid recomputation. Recently, frontier open-source LLMs adopt sparse attention to reduce computation by selecting only part of the history, while SSDs provide a cheaper alternative to CPU DRAM for storing KV caches. However, sparse KV selection depends on the ad hoc intermediate values during model inference, so it forces SSD reads to lie on the inference critical path. These reads are further slowed by fragmented accesses and read-write interference in SSDs. To address these challenges, we present Janus, an agentic serving framework for sparse attention LLMs with SSD-centric KV storage. Janus focuses on append prefill, which processes each round's newly added inputs and accounts for most history KV loading. To move SSD reads out of the critical path, Janus runs the model's own KV selection module on earlier intermediate values, predicting KV demand without additional training. The predicted reads overlap with model computation, and any prediction misses are fetched before attention executes to preserve model outputs. To improve SSD efficiency, Janus coalesces adjacent reads, packs scattered KV pages into sequential writes on the CPU, and limits background writes while reads are active. Across three models and three agentic traces, Janus outperforms existing works by up to 1.57-3.69 times (1.22-1.85 times on average) in terms of the time to first token latency, while maintaining decode efficiency.

为什么值得关注

KV 分层从 GPU-DRAM 一路伸到 SSD;稀疏注意力给了预测失误时的补救空间。未提开源,与 TempoKV、DPS 同方向留观。

English Summary

Janus serves sparse-attention LLM agents with SSD-backed KV storage. It targets append prefill — the bulk of history KV loading — by running the model's own KV-selection module on earlier intermediate values to predict demand (training-free), overlapping predicted reads with computation and fetching misses before attention executes so outputs are unchanged; SSD-side read coalescing, CPU-side packing of fragmented pages into sequential writes, and read-priority scheduling improve device efficiency. Across three models and three agentic traces, TTFT improves by up to 1.57-3.69x over prior work (1.22-1.85x on average) with decode efficiency maintained.