基础设施 3.0 · 值得看 2026-09-28 · 论文

Spexis: Speculative Lookahead Scheduling for LLM Inference

把投机执行从省单请求时延的手法升级为新的并行轴:Spexis 让 speculation 与正常执行并行跑,叠在流水线并行与张量并行之上不增加 KV cache 内存,从而改善多卡推理的内存效率与瓶颈;lookahead 调度预测投机质量与未来内存压力,减少白做的投机KV 驱逐与重算基于 vLLM 实现,多种 GPU 配置下相对最优 PP+TP 组合基线最高加速 34%,代码开源EMNLP 2026 main

打开原文回到归档

Spexis: Speculative Lookahead Scheduling for LLM Inference

  • ID: 05d2de45
  • 原文链接: https://arxiv.org/abs/2609.34370
  • 作者: Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee, Jiwon Seo
  • 日期: 2026-09-28
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: speculative-decoding, llm-serving, parallelism, vllm
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

把投机执行从省单请求时延的手法升级为新的并行轴:Spexis 让 speculation 与正常执行并行跑,叠在流水线并行与张量并行之上、不增加 KV cache 内存,从而改善多卡推理的内存效率与瓶颈;lookahead 调度预测投机质量与未来内存压力,减少白做的投机、KV 驱逐与重算。基于 vLLM 实现,多种 GPU 配置下相对「最优 PP+TP 组合」基线最高加速 34%,代码开源。EMNLP 2026 main。

论文信息

  • arXiv ID: 2609.34370
  • 提交日期: 2026-09-28
  • arXiv 分类: cs.LG, cs.DC
  • 作者: Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee, Jiwon Seo
  • 链接: https://arxiv.org/abs/2609.34370

Abstract(arXiv 原文)

Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.

为什么值得关注

多卡 serving 的并行设计空间里多了一个成本不高的维度:投机作为并行轴不占 KV 内存,vLLM 上的实现可以直接试。

English Summary

Spexis treats speculative execution as a new parallelism axis alongside pipeline and tensor parallelism, running speculation concurrently with normal execution without increasing KV-cache memory, and uses lookahead scheduling to predict speculation quality and future memory pressure to cut wasted speculation, eviction and recomputation. Built on vLLM, it speeds serving by up to 34% over the optimal PP+TP baseline across GPU configurations. EMNLP 2026 main; code open source.