基础设施 4.0 · 优秀 2026-08-11 · 论文

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

用户在边缘节点间切换时,KV cache 要在移动性约束的时间窗内传完,并发切换会打满回传带宽ImpactHO 按重要性对每个用户的 cache 排序只传信息量最大的那部分,把传输建模为多用户回传分配问题,最优解是闭式加权注水推广了信息论注水且支持在线调度500ms 传输窗内平均精度 93.7%+(距全 cache 上界 0.5pp 以内),达到千里眼上界的 98.2-99.5%

打开原文回到归档

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

  • ID: d101b03b
  • 原文链接: https://arxiv.org/abs/2608.10545
  • PDF: https://arxiv.org/pdf/2608.10545v1
  • 作者: Minwoo Kim, Soochang Song, Namyoon Lee, Bang Chul Jung, Yongjune Kim
  • 日期: 2026-08-11
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, edge-computing, handover, water-filling, bandwidth-allocation
  • 质量评分: 4/5
  • 抓取时间: 2026-08-17T23:51:34+08:00

中文导读

用户在边缘节点间切换时,KV cache 要在移动性约束的时间窗内传完,并发切换会打满回传带宽。ImpactHO 按重要性对每个用户的 cache 排序、只传信息量最大的那部分,把传输建模为多用户回传分配问题,最优解是闭式加权注水——推广了信息论注水且支持在线调度。500ms 传输窗内平均精度 93.7%+(距全 cache 上界 0.5pp 以内),达到千里眼上界的 98.2-99.5%。

为什么值得关注

边缘切换的 KV 传输按重要性排序 + 加权注水分配:500ms 窗内精度贴住全量上界。

收录理由:KV cache 压缩三部曲之三(迁移带宽轴),把传输调度变成有闭式解的凸分配问题

Abstract

Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node. However, simultaneous handovers saturate the backhaul, preventing full cache delivery within the mobility-imposed transfer window. Rather than allocating bandwidth as if all cache entries were equally valuable, we order each user's KV cache by importance and transmit only its most informative fraction, turning token-level sparsity into communication savings. We cast the transfer as a multi-user backhaul allocation problem that maximizes average accuracy across users. Each user's partial-cache accuracy serves as its utility: a sigmoid that fits measurements on the RULER benchmark with $R^2>0.99$ across models and context lengths. Because importance ordering front-loads the high-value entries, the concave region of the accuracy curve spans nearly the entire cache. Our proposed allocator keeps served users within this region, making each per-slot allocation problem convex. The optimum is derived via a closed-form weighted water-filling solution that generalizes information-theoretic water-filling and enables online scheduling. The proposed allocator attains over 93.7% average accuracy in a 500ms transfer window, within 0.5pp of the full-cache ceiling, and reaches 98.2-99.5% of a clairvoyant upper bound.

元数据

  • arXiv ID: 2608.10545
  • 主分类: cs.NI
  • 分类: cs.NI, cs.AI, cs.DC
  • 评论: N/A
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.10545(2026-08-17)。