KV-streams for Efficient Compaction in Agentic Reinforcement Learning
- ID: e03344a9
- 原文链接: https://arxiv.org/abs/2609.35750
- PDF: https://arxiv.org/pdf/2609.35750v2
- 作者: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie
- 日期: 2026-09-28
- 更新: 2026-09-29
- 分类: infra
- 来源类型: paper
- 标签: agentic-rl, kv-cache, context-compaction, training-throughput
- 质量评分: 4/5
- 抓取时间: 2026-09-30T12:24:18Z
中文导读
拉长智能体 RL 的任务时程被 GPU 显存中的长轨迹上下文卡住,主流上下文压缩策略每次压缩都要反复 prefill,拖垮训练吞吐KV-streams 提出流式前传 KV cache 而非压缩后冲刷,与任意压缩策略即插即用:在三种压缩策略上取得 2.65 倍训练 wall-clock 加速且未见性能受损额外发现:流式保留的 KV cache 可充当循环状态,携带早已离开上下文窗口的信息;受控实验显示仅靠 RL 即可涌现这一行为,与此前需要额外机制的结论相反定位为任何 post-training 流水线的轻量即插即用组件
为什么值得关注
拉长智能体 RL 的任务时程被 GPU 显存中的长轨迹上下文卡住,主流上下文压缩策略每次压缩都要反复 prefill,拖垮训练吞吐KV-streams 提出流式前传 KV cache 而非压缩后冲刷,与任意压缩策略即插即用:在三种压缩策略上取得 2.
- arXiv categories: c, s, ., L, G, ,, , c, s, ., A, I
- published 2026-09-28, updated 2026-09-29; 18 authors
- fit: entry category
infra, tags agentic-rl, kv-cache, context-compaction, training-throughput
关键信息
- 论文标题: KV-streams for Efficient Compaction in Agentic Reinforcement Learning
- 作者: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie
- arXiv: https://arxiv.org/abs/2609.35750
- 发布时间: 2026-09-28
- arXiv 分类: c, s, ., L, G, ,, , c, s, ., A, I
- 关联标签: agentic-rl, kv-cache, context-compaction, training-throughput
English Abstract
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
English Summary
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training....
Obsidian Notes
- 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
- 中文导读与价值判断锚定在条目摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。