Agent 与自动化 4.0 · 优秀 2026-09-22 · 论文

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

LLM agent 处理相关任务流时,反复在上下文内重建同样的控制决策Growing Harness 提出失败引导的 harness 训练范式:从只暴露固定模型/工具接口不含任务控制器的 strategy-free scaffold 出发,用函数级执行轨迹把每次失败定位到有界代码面,优化器联合修复一批失败,成功优先的 held-out 门控回滚损害既有能力的修复在 BrowseComp-Plus 与 WebArena-Verified 上跨 4B-120B 三种模型,相比 Tool-Calling agent 减少 76.0-91.8% 的 LLM 调用与 74.4-98.6% 推理成本;4B 模型上成功率保持 44.7%+,而 Tool-Calling 跌到 6.7%控制结构从任务反馈中长进代码而非塞进上下文

打开原文回到归档

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

  • ID: 76147fee
  • 原文链接: https://arxiv.org/abs/2609.26760
  • PDF: https://arxiv.org/pdf/2609.26760v1
  • 作者: Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao
  • 发布日期: 2026-09-22
  • 更新日期: 2026-09-22
  • 分类: cs.AI
  • 来源类型: paper
  • 标签: arxiv, agent-harness, scaffold, context-engineering, paper
  • 质量评分: 4/5
  • 抓取时间: 2026-09-24T04:23:22Z

中文导读

LLM agent 处理相关任务流时,反复在上下文内重建同样的控制决策Growing Harness 提出失败引导的 harness 训练范式:从只暴露固定模型/工具接口不含任务控制器的 strategy-free scaffold 出发,用函数级执行轨迹把每次失败定位到有界代码面,优化器联合修复一批失败,成功优先的 held-out 门控回滚损害既有能力的修复在 BrowseComp-Plus 与 WebArena-Verified 上跨 4B-120B 三种模型,相比 Tool-Calling agent 减少 76.0-91.8% 的 LLM 调用与 74.4-98.6% 推理成本;4B 模型上成功率保持 44.7%+,而 Tool-Calling 跌到 6.7%控制结构从任务反馈中长进代码而非塞进上下文

为什么值得关注

让 harness 从失败轨迹里自己长出来:LLM 调用省 76-92%,4B 小模型成功率 44.7% 而 tool-calling 只有 6.7%

关键信息

  • 论文标题:Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
  • 作者:Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao
  • arXiv:https://arxiv.org/abs/2609.26760
  • 发布时间:2026-09-22
  • 更新时间:2026-09-22
  • arXiv 分类:cs.AI, cs.SE
  • 可选注释:16 pages, 6 figures
  • 关联标签:arxiv, agent-harness, scaffold, context-engineering, paper

English Abstract

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.

English Summary

LLM agents handling streams of related tasks repeatedly reconstruct the same control decisions inside each task's context. Growing Harness is a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold exposing fixed model and tool interfaces but no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs windows of failures jointly, and a success-first held-out gate rolls back repairs that harm prior capability. Across BrowseComp-Plus and WebArena-Verified with models from 4B to 120B, it reduces LLM calls by 76.0-91.8% and inference cost by 74.4-98.6% versus a Tool-Calling agent, and keeps 44.7-45.3% success across model scales where Tool-Calling falls to 6.7% at 4B.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页面属于 content-fetcher 能动下的高分筻补东生成,不修改原条目字段,仅产出可读的中英双语页面。