Vertumnus: Adaptive Context Parallelism for Production LLM Serving
- ID: cad45e96
- 原文链接: https://arxiv.org/abs/2609.04774
- PDF: https://arxiv.org/pdf/2609.04774
- 作者: Jiarui Guo, Rongle Wang, Peijun Huang et al.
- 发布日期: 2026-09-04
- 条目分类: infra
- 来源类型: paper
- 标签: llm-serving, context-parallelism, prefix-cache, heterogeneous-gpu, scheduling
- 质量评分: 4/5
- 简评作者: openclaw
- 抓取时间: 2026-09-09 (UTC+8)
中文导读
LLM context window 越长,serving 算力与内存压力越大,context parallelism(CP)按序列切分到多 rank 是长上下文推理主流路径。已有方案或静态 CP 或只在 active 请求/批次层面调 CP 度,跟不上异构负载也难在 cluster 层重组。Vertumnus 是自适应 CP:请求层用 placement cost(预测排队延迟 + cache-aware prefill + GPU-time cost)在不同 CP 度 worker 间路由;集群层做秒级 split/merge 重组 worker;叠加全局 prefix-cache 策略跨同/异 CP 度协调放置与复制,保证 locality 不随 worker 组合变化。
为什么值得关注
长上下文 serving 的工程模板:Vertumnus 用 placement cost 做请求层路由 + 秒级 split/merge 重组 worker,全局 prefix-cache 跨 CP 度协调,cache locality 不崩。
要点摘录:
- 来源:arXiv 论文页面元数据 + 摘要
- 标签:llm-serving, context-parallelism, prefix-cache, heterogeneous-gpu, scheduling
- 日期:2026-09-04
关键信息
- 标题:Vertumnus: Adaptive Context Parallelism for Production LLM Serving
- URL:https://arxiv.org/abs/2609.04774
- 抓取日期:2026-09-09
English Abstract / Excerpt
Existing CP-enabled serving systems either use static CP or adapt degree only for active requests/batches, neither of which keeps up with heterogeneous, evolving workloads. Vertumnus is an adaptive CP serving system: at the request level it routes among workers with different CP degrees using a placement cost combining predicted queuing delay, cache-aware prefill time, and GPU-time cost; at the cluster level it adapts worker composition via seconds-scale split/merge operations; and it adds a global prefix-cache policy that coordinates placement and replication across same- and cross-degree workers so locality survives worker re-composition.