基础设施 4.0 · 优秀 2026-08-13 · 论文

vToken: Token-Level Virtualization for Reclaimable KV Caches

PagedAttention 的定长块消掉了分配器级碎片,但 token 粒度的 KV 驱逐比块管理更细,块内碎片大量不可回收vToken 加一层轻量 token 级虚拟化:用 token 表间接层把逻辑存活与物理摆放解耦,维持稳定逻辑视图,异步 repack 存活 token 实现物理回收;保持 PagedAttention kernel 与 CUDA Graph 兼容对比基线每请求保留 KV 块降 27.2%-72.3%,SLA 约束吞吐最高提升 1.37x

打开原文回到归档

vToken: Token-Level Virtualization for Reclaimable KV Caches

  • ID: d8f43cd5
  • 原文链接: https://arxiv.org/abs/2608.13263
  • PDF: https://arxiv.org/pdf/2608.13263v1
  • 作者: Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li
  • 日期: 2026-08-13
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, paged-attention, virtualization, memory-fragmentation, serving
  • 质量评分: 4/5
  • 抓取时间: 2026-08-17T23:51:34+08:00

中文导读

PagedAttention 的定长块消掉了分配器级碎片,但 token 粒度的 KV 驱逐比块管理更细,块内碎片大量不可回收。vToken 加一层轻量 token 级虚拟化:用 token 表间接层把“逻辑存活”与“物理摆放”解耦,维持稳定逻辑视图,异步 repack 存活 token 实现物理回收;保持 PagedAttention kernel 与 CUDA Graph 兼容。对比基线每请求保留 KV 块降 27.2%-72.3%,SLA 约束吞吐最高提升 1.37x。

为什么值得关注

KV cache 的 token 级虚拟化:块内碎片可回收,每请求保留块最高降七成。

收录理由:KV cache 压缩三部曲之二(回收粒度轴),用间接层解耦逻辑存活与物理布局

Abstract

Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable. We present vToken, a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement. vToken maintains a stable logical token view through token-table indirection and realizes physical reclamation by repacking live tokens asynchronously. The design preserves PagedAttention kernels and CUDA Graph compatibility. We implement vToken in vLLM and evaluate it with H2O, Random, and Scissorhands across models. Compared with a paired Naive-Evict baseline, vToken reduces retained KV blocks per request by 27.2\%--72.3\% and improves SLA-constrained throughput by up to 1.37$\times$. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2$\times$, while reducing the per-policy integration footprint from 500+ lines to under 50.

元数据

  • arXiv ID: 2608.13263
  • 主分类: cs.AI
  • 分类: cs.AI, cs.DC, cs.OS
  • 评论: N/A
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.13263(2026-08-17)。