AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
- ID: 1176bd48
- 原文链接: https://arxiv.org/abs/2610.02163
- PDF: https://arxiv.org/pdf/2610.02163
- 作者: Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong
- 日期: 2026-10-01
- 更新: 2026-10-01
- 分类: agents
- 来源类型: paper
- 标签: context-management, compaction, coding-agents, reinforcement-learning, swe-bench
- 质量评分: 4/5
- Fetch: 2026-10-03T04:21:07Z
中文导读
AutoCompact 把何时压缩上下文保留什么状态如何从压缩点继续训练成编码 agent 策略的一部分:先让基础 agent 跑仓库级任务,用 judge 审查其压缩决策/摘要/后续动作,有缺陷的输出在环境中执行前替换为修正版本,轨迹从修正后的决策继续;再用这些轨迹做 SFT,随后以任务成功率为奖励做编码与压缩的联合 RLSWE-bench Verified 绝对提升 9.2%,SWE-PolyBench Verified 提升 5.0%,且在 256K 窗口从不溢出16K 窗口触发回退压缩的设置下都成立核心主张是上下文压缩不应是机械的溢出处理,而是可学习的策略行为
为什么值得关注
AutoCompact 把上下文压缩从溢出应急变成可学习策略:judge 修正轨迹 + 联合 RL,SWE-bench 绝对 +9.2%
以上导读与价值判断锚定论文摘要与元数据,完整英文摘要见下文。
关键信息
- 论文标题:AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
- 作者:Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong
- arXiv:https://arxiv.org/abs/2610.02163
- 发布时间:2026-10-01
- arXiv 分类:cs.CL
- 关联标签:context-management, compaction, coding-agents, reinforcement-learning, swe-bench
English Abstract
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
English Summary
AutoCompact trains a coding agent to make context-compaction decisions (when to compact, what working state to preserve, how to continue) as part of its policy. Training data: run the base agent on coding tasks, use a judge to review compaction decisions/summaries/post-compaction actions, replace flawed outputs with corrected ones before environment execution so each trajectory continues from corrected decisions; SFT on these trajectories, then joint RL on coding and compaction with task-success rewards. Absolute pass-rate gains of 9.2% on SWE-bench Verified and 5.0% on SWE-PolyBench Verified, holding across inference budgets with a never-overflowing 256K window and a 16K window with fallback compaction.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。