基础设施 4.0 · 优秀 2026-09-04 · 论文

Vertumnus: Adaptive Context Parallelism for Production LLM Serving

LLM context window 越长,serving 算力与内存压力越大,context parallelism(CP)按序列切分到多 rank 是长上下文推理主流路径已有方案或静态 CP 或只在 active 请求/批次层面调 CP 度,跟不上异构负载也难在 cluster 层重组Vertumnus 是自适应 CP:请求层用 placement cost(预测排队延迟 + cache-aware prefill + GPU-time cost)在不同 CP 度 worker 间路由;集群层做秒级 split/merge 重组 worker;叠加全局 prefix-cache 策略跨同/异 CP 度协调放置与复制,保证 locality 不随 worker 组合变化

打开原文回到归档

Vertumnus: Adaptive Context Parallelism for Production LLM Serving

  • ID: cad45e96
  • 原文链接: https://arxiv.org/abs/2609.04774
  • PDF: https://arxiv.org/pdf/2609.04774
  • 作者: Jiarui Guo, Rongle Wang, Peijun Huang et al.
  • 发布日期: 2026-09-04
  • 条目分类: infra
  • 来源类型: paper
  • 标签: llm-serving, context-parallelism, prefix-cache, heterogeneous-gpu, scheduling
  • 质量评分: 4/5
  • 简评作者: openclaw
  • 抓取时间: 2026-09-09 (UTC+8)

中文导读

LLM context window 越长,serving 算力与内存压力越大,context parallelism(CP)按序列切分到多 rank 是长上下文推理主流路径。已有方案或静态 CP 或只在 active 请求/批次层面调 CP 度,跟不上异构负载也难在 cluster 层重组。Vertumnus 是自适应 CP:请求层用 placement cost(预测排队延迟 + cache-aware prefill + GPU-time cost)在不同 CP 度 worker 间路由;集群层做秒级 split/merge 重组 worker;叠加全局 prefix-cache 策略跨同/异 CP 度协调放置与复制,保证 locality 不随 worker 组合变化。

为什么值得关注

长上下文 serving 的工程模板:Vertumnus 用 placement cost 做请求层路由 + 秒级 split/merge 重组 worker,全局 prefix-cache 跨 CP 度协调,cache locality 不崩。

要点摘录:

  • 来源:arXiv 论文页面元数据 + 摘要
  • 标签:llm-serving, context-parallelism, prefix-cache, heterogeneous-gpu, scheduling
  • 日期:2026-09-04

关键信息

English Abstract / Excerpt

Existing CP-enabled serving systems either use static CP or adapt degree only for active requests/batches, neither of which keeps up with heterogeneous, evolving workloads. Vertumnus is an adaptive CP serving system: at the request level it routes among workers with different CP degrees using a placement cost combining predicted queuing delay, cache-aware prefill time, and GPU-time cost; at the cluster level it adapts worker composition via seconds-scale split/merge operations; and it adds a global prefix-cache policy that coordinates placement and replication across same- and cross-degree workers so locality survives worker re-composition.