An Empirical Study of Harness Design for Coding Agents
- ID: 3f8ce213
- 原文链接: https://arxiv.org/abs/2609.20804
- PDF: https://arxiv.org/pdf/2609.20804v1
- 作者: Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
- 日期: 2026-09-17
- 更新: 2026-09-17
- 分类: agents
- 来源类型: paper
- 标签: coding-agents, harness-design, context-management, empirical-study, swe-bench
- 质量评分: 4/5
- 抓取时间: 2026-09-20T04:24:27Z
中文导读
harness 设计实证研究:固定执行循环只动 planning/action space/context management 三个组件,176 组匹配配置4 个模型SWE-Bench Verified 与 Terminal-Bench 2.1结论直接可用于 coding agent 搭建:上下文管理在窗口预算紧张时价值最大,收益主要来自防止 context-overflow 失败;规则裁剪放在 LLM 摘要之前效率最好,让被裁剪内容可恢复反而几乎没人用摘要信息完整,可 grounded 入库
为什么值得关注
harness 设计实证研究:固定执行循环只动 planning/action space/context management 三个组件,176 组匹配配置4 个模型SWE-Bench Verified 与 Terminal-Bench 2.
关键信息
- 论文标题:An Empirical Study of Harness Design for Coding Agents
- 作者:Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
- arXiv:https://arxiv.org/abs/2609.20804
- 发布时间:2026-09-17
- arXiv 分类:cs.AI, cs.CL, cs.LG, cs.SE
- 关联标签:coding-agents, harness-design, context-management, empirical-study, swe-bench
- 备注:43 pages
English Abstract
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.
English Summary
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读、价值判断、关键事实均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- 抓取时间:2026-09-20T04:24:27Z
- 抓取来源:opencli arxiv paper 2609.20804