基础设施 4.0 · 优秀 2026-09-28 · 论文

Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs

面向移动 NPU 的 LLM serving KV 复用系统设计:现有 KV 缓存为云 GPU 动态执行环境优化,而移动 NPU 计算图需静态编译内存容量与 I/O 带宽双受限方案是计算-存储协同设计图内机制把选择性 KV 重计算映射进静态 NPU 图,调和算法动态性与 NPU 静态性;图间调度器用动态规划做 chunk 合并与 padding 最小化;再加 tree-hash-semantic 混合层级 KV 管理器成本感知预取/驱逐,以及把 KV 加载rerotation存储与 NPU 执行重叠的二维流水线代表性端侧负载上 TTFT 较无复用与仅前缀缓存降 40-60%EuroSys 2027 Spring Cycle 接收

打开原文回到归档

Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs

  • ID: 195ec4b4
  • 原文链接: https://arxiv.org/abs/2609.34727
  • 作者: Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun, Zhaode Wang, Zeyu Zhao, Chengfei Lv, Fan Wu, Guihai Chen
  • 日期: 2026-09-28
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, npu, on-device, llm-serving
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

面向移动 NPU 的 LLM serving KV 复用系统设计:现有 KV 缓存为云 GPU 动态执行环境优化,而移动 NPU 计算图需静态编译、内存容量与 I/O 带宽双受限。方案是计算-存储协同设计——图内机制把选择性 KV 重计算映射进静态 NPU 图,调和算法动态性与 NPU 静态性;图间调度器用动态规划做 chunk 合并与 padding 最小化;再加 tree-hash-semantic 混合层级 KV 管理器、成本感知预取/驱逐,以及把 KV 加载、rerotation、存储与 NPU 执行重叠的二维流水线。代表性端侧负载上 TTFT 较无复用与仅前缀缓存降 40-60%。EuroSys 2027 Spring Cycle 接收。

论文信息

  • arXiv ID: 2609.34727
  • 提交日期: 2026-09-28
  • arXiv 分类: cs.OS, cs.AI, cs.DB
  • 作者: Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun, Zhaode Wang, Zeyu Zhao, Chengfei Lv, Fan Wu, Guihai Chen
  • 链接: https://arxiv.org/abs/2609.34727

Abstract(arXiv 原文)

On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily optimized for cloud GPUs with dynamic execution environments and abundant memory bandwidth. These architectural assumptions do not hold on mobile NPUs, where computation graphs must be statically compiled and both memory capacity and I/O bandwidth are severely constrained. In this work, we present a compute-storage co-design for mobile-centric prefix and non-prefix KV reuse. We first propose an intra-graph mechanism that maps selective KV recomputation onto static NPU graphs, reconciling algorithmic dynamicity with NPU staticity. We further develop an inter-graph scheduler to optimize chunk merging and minimize padding with dynamic programming. To address mobile bandwidth limitations, we introduce a hierarchical KV manager featuring a tree-hash-semantic hybrid structure, along with cost-aware prefetching and eviction policies. We also build a two-dimensional pipeline that overlaps KV loading, rerotation, and storage with NPU execution, hiding data-movement latency. Experiments across representative on-device workloads and LLMs show that our design reduces time-to-first-token (TTFT) by $40-60\%$ compared with no reuse and prefix-only caching.

为什么值得关注

端侧长上下文推理的瓶颈正在从算子执行转向 KV 数据面。这篇把「算法动态性 vs NPU 静态性」的矛盾拆成图内/图间/存储三层来解,是本周端侧方向首读。

English Summary

A compute-storage co-design for KV reuse on mobile NPUs, where statically compiled graphs and tight memory/I-O budgets break cloud-GPU caching assumptions. An intra-graph mechanism maps selective KV recomputation onto static NPU graphs; an inter-graph scheduler uses dynamic programming for chunk merging and padding minimization; a hierarchical tree-hash-semantic KV manager with cost-aware prefetching/eviction and a two-dimensional pipeline overlapping KV movement with NPU execution cut TTFT by 40-60% versus no reuse and prefix-only caching. Accepted to EuroSys 2027 (Spring cycle).