The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
- ID: 49e9d662
- 原文链接: https://arxiv.org/abs/2608.16630
- PDF: https://arxiv.org/pdf/2608.16630
- 作者: Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler
- 日期: 2026-08-17
- 更新: 2026-08-17
- 分类: cs.SE, cs.AI, cs.LG
- 来源类型: arxiv
- 标签: coding-agent, context-engineering, benchmark, arxiv, paper
- 质量评分: 5/5
- 抓取时间: 2026-08-20T15:44:49Z
中文导读
论文把仓库级编码任务建模为耦合事实图的重建:每次编辑所需事实来自近期上下文或参数记忆,两边都覆盖不到的部分构成 coherence debt(一致性债)。在 7 个模型、5 种 harness 上做事实供给/扣留与故障注入:两个通道全空时没有模型能完成陌生 API 任务;对真实库重命名后 7 个模型全部在同一位置失败。决定结果的是事实可得性而非距离,不同 harness 达到同样通过率的 token 消耗相差超 10 倍;事实被扣留时多花 token 什么都换不回,且缺事实时 agent 倾向产出错误工作(编造文件、猜测取值)而非停滞。
为什么值得关注
用耦合事实图刻画 coding agent 的上下文策略:事实可得性而非距离决定成败,harness 间 token 效率相差 10 倍
English Abstract
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.
Obsidian 证据
- 来源 digest: 论文流水线 2026-08-20(评分 8.7)。
- 元数据与摘要经 opencli arxiv paper 核对;中文导读锚定摘要陈述的事实与数字。