Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting
- source_url: https://arxiv.org/abs/2607.00038
- source_type: paper
- platform: arxiv
- author: Sandeco Macedo
- original_date: 2026-06-28
- added_date: 2026-07-22
- category: coding
- tags: loop-engineering, coding-agents, verification, harness, arxiv
- quality_score: 5
- arxiv_id: 2607.00038
- arxiv_categories: cs.SE
- pdf_url: https://arxiv.org/pdf/2607.00038v1
摘要(中文)
把 2026 中期口号「别再一步步喂 prompt,设计会自己跑的 loop」正式化为 loop specification:触发、目标、验证、停止规则、记忆五件套,可交给 Claude Code/Codex 等 harness 自主推进。区分外部 loop 规格、普通程序循环与 harness 内置感知-行动环;认为 loop engineering 是 prompt→context→harness→loop 递进的新层,但并不取代 prompt engineering。手工编码 50 个公开 loop 语料:约 70% 验证落在自主区、74% 命名终端态,自动触发与持久记忆仍弱。并给出自纠错、reward hacking、model-as-judge 脆弱性相关的设计原则与反模式。
Summary (English)
Treats the mid-2026 slogan—stop step-by-step prompting, design the loop that prompts the agent—as a serious object of study. Defines the loop specification: a bounded reusable artifact of trigger, goal, verification, stopping rule, and memory handed to a harness (e.g. Claude Code or Codex). Distinguishes external loop specs from ordinary programming loops and the harness's internal perceive-act cycle. Positions loop engineering as a layer after prompt/context/harness without retiring prompt engineering. Codes a public corpus of fifty real loops: ~70% verify in the autonomous zone and ~74% name terminal states, while automated triggering and durable memory remain underdeveloped. Offers design principles and anti-patterns grounded in self-correction, reward hacking, and model-as-judge fragility.
One-liner
Coding agent 下一层是 loop specification:触发、目标、验证、停止规则与记忆。
Source body / metadata
arXiv abstract grounded intake for 2607.00038. PDF: https://arxiv.org/pdf/2607.00038v1
Treats the mid-2026 slogan—stop step-by-step prompting, design the loop that prompts the agent—as a serious object of study. Defines the loop specification: a bounded reusable artifact of trigger, goal, verification, stopping rule, and memory handed to a harness (e.g. Claude Code or Codex). Distinguishes external loop specs from ordinary programming loops and the harness's internal perceive-act cycle. Positions loop engineering as a layer after prompt/context/harness without retiring prompt engineering. Codes a public corpus of fifty real loops: ~70% verify in the autonomous zone and ~74% name terminal states, while automated triggering and durable memory remain underdeveloped. Offers design principles and anti-patterns grounded in self-correction, reward hacking, and model-as-judge fragility.
Obsidian evidence
- local_note: OpenClaw定时任务/研究写作流水线/2026-07-22-研究写作流水线.md
- intake_run: daily-intake-evening 2026-07-22