模型与实验室 4.0 · 优秀 2026-08-26 · 论文

Prefix Sliding for efficient test-time scaling

发现推理过程中大部分中间 token 随推理推进失去重要性,据此提出 Prefix Sliding:推理时丢弃既不属于前缀(关键指令与工具)也不属于最近几千 token 窗口的中间 token,把内存上限与推理长度解耦,实现高效长程测试时扩展免训练即可让现有模型提速 3 倍且性能不掉;配合强化学习训练则能把推理轨迹扩展到 10 万 token 以上并取得更好成绩,消融显示优于摘要中间 token 与普通滑窗

打开原文回到归档

Prefix Sliding for efficient test-time scaling

  • ID: cf678d0e
  • 原文链接: https://arxiv.org/abs/2608.26070
  • PDF: https://arxiv.org/pdf/2608.26070v1
  • 作者: Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt, Percy Liang, Jason Wei, Andrew Y. Ng, Luke Zettlemoyer, Yejin Choi, Mike Lewis
  • 日期: 2026-08-26
  • 更新: 2026-08-26
  • 分类: models
  • 来源类型: paper
  • 标签: test-time-scaling、efficiency、long-context、reasoning、arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-28T05:15:03Z

中文导读

发现推理过程中大部分中间 token 随推理推进失去重要性,据此提出 Prefix Sliding:推理时丢弃既不属于前缀(关键指令与工具)也不属于最近几千 token 窗口的中间 token,把内存上限与推理长度解耦,实现高效长程测试时扩展免训练即可让现有模型提速 3 倍且性能不掉;配合强化学习训练则能把推理轨迹扩展到 10 万 token 以上并取得更好成绩,消融显示优于摘要中间 token 与普通滑窗

为什么值得关注

核心发现是大部分中间推理 token 随推理推进失去重要性:丢掉前缀与最近窗口之外的 token,免训练 3 倍提速性能不掉,RL 训练后轨迹可扩展到 10 万 token 以上且成绩更好;对做长程推理 serving 与 test-time scaling 的工程团队是直接可用的杠杆。

关键信息

  • 论文标题:Prefix Sliding for efficient test-time scaling
  • 作者:Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt, Percy Liang, Jason Wei, Andrew Y. Ng, Luke Zettlemoyer, Yejin Choi, Mike Lewis
  • arXiv:https://arxiv.org/abs/2608.26070
  • 发布时间:2026-08-26
  • arXiv 分类:cs.CL, cs.AI, cs.LG
  • 关联标签:test-time-scaling、efficiency、long-context、reasoning、arxiv

English Abstract

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens. The prefix has key instructions and tools available to the model, while the most recent tokens are the current reasoning the model is working on. This caps the total memory requirement regardless of how long the model reasons, allowing for efficient long-horizon test-time scaling. Without training, Prefix Sliding can make existing models 3x faster while maintaining performance. Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens. Ablations show Prefix Sliding outperforms summarizing intermediate tokens or vanilla sliding window. Our code is at https://github.com/Muennighoff/prefix-sliding

English Summary

Finds that most intermediate reasoning tokens lose importance as a model continues reasoning, and proposes Prefix Sliding: during reasoning, discard tokens outside the prefix (key instructions and tools) and the window of the last few thousand tokens. This caps total memory regardless of reasoning length, enabling efficient long-horizon test-time scaling. Without training it makes existing models 3x faster while maintaining performance; training with Prefix Sliding via reinforcement learning scales to reasoning traces beyond a hundred thousand tokens and achieves better performance. Ablations show it outperforms summarizing intermediate tokens or vanilla sliding windows.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。