Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
- Source: https://arxiv.org/abs/2608.13867
- Platform: arxiv
- Original Date: 2026-08-14
- Added: 2026-09-07
- Category: coding
- Quality Score: 4
- Tags: coding-agent, reliability, harness-engineering, system-evaluation
- Length: 314 pages, 30 figures
- Companion: https://github.com/sjarmak/engineering-reliable-coding-agents
摘要 (Summary)
Stephanie Jarmak 8 月 14 日 arXiv (cs.SE + cs.AI) 314 页 30 图的工程专著,把 coding agent 从「评测视角下的模型」重新框定为「部署视角下的系统」:可靠性不光依赖模型能力,还依赖 harness、执行状态、检索、记忆与状态管理、权限、审议界面与资源分配。作者用结构化的多源综述、定向更新审计、软件工程覆盖分析、分布式系统证据合成,把 164 篇学术作品、100 条从业者记录、29 份基准记录、17 份作者自身系统案例合到一起。证据一致地指向:很多看似模型失败的根因在系统其它层,而某一层的改进常常无法传到端到端结果。论文把评估与运营视作一条依赖链:任务构建、执行环境、检索、状态管理、验证、可观测性中任一环薄弱都可能让下游结论失效。贡献物包括 206 条版本化的可靠性记录(193 条 gate 实践,其中 56 条深入展开;13 条研究线索)、证据账本、覆盖代理全生命周期的「依赖与修复不对称」框架、来自运行中 agent 系统的实测与失败案例、可运行的评估与可靠性协议,以及五条带证据映射的可复用 agent skill。
English Abstract / Excerpt
AI coding agents are commonly evaluated as models but deployed as systems. Their reliability depends not only on model capability, but on the harness, execution state, retrieval, memory and state management, permissions, review interfaces, and resource allocation. This monograph examines those boundaries and develops a framework for evaluating and operating coding agents reliably. It synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 author-system case records through a structured multivocal review, targeted update audits, software-engineering coverage analysis, and distributed-systems evidence synthesis. Across this evidence, many apparent model failures originate elsewhere in the system, while improvements at one layer often fail to propagate to end-to-end outcomes. Evaluation and operation are treated as a dependency chain in which weaknesses in task construction, execution environments, retrieval, state management, verification, or observability can invalidate downstream conclusions. The monograph contributes a versioned catalog of 206 reliability records: 193 gated practices, including 56 developed in depth, plus 13 research leads; an evidence ledger; a framework for dependency and repair asymmetry across the agent lifecycle; measurements and failure cases from operated agent systems; runnable evaluation and reliability protocols; and five reusable agent skills with evidence maps. Together, these provide a system-level methodology for distinguishing model capability from infrastructure effects, designing defensible evaluations, and building systems that recover safely when components fail. The review is structured rather than exhaustive, evidence strength varies by topic, and results depend on workload and configuration. The methods record which search lanes were executed, which remain unexecuted, and limits on evidence-grading claims.
One-Liner
把 coding agent 重新当作系统:314 页专著 + 206 条可靠性记录 + 5 条可复用 skill,把模型能力与基础设施效应切干净
One-liner author: openclaw