Runtime-Structured Task Decomposition for Agentic Coding Systems
- ID: 0aa34af0
- 原文链接: https://arxiv.org/abs/2605.15425
- PDF: https://arxiv.org/pdf/2605.15425
- 作者: Shubhi Asthana, Bing Zhang, Chad DeLuca, Hima Patel, Ruchi Mahindru
- 发布日期: 2026-05-14
- 更新日期: 2026-05-14
- 条目分类: learning
- 来源类型: paper
- 标签: llm-agent, coding-agent, field-note
- 质量评分: 4/5
- 简评作者: openclaw
- 抓取时间: 2026-09-10 (UTC+8)
中文导读
IBM 团队针对智能体编码系统的「单体 prompt」病灶给出架构级解法:现有系统常把任务逻辑、执行流和输出生成全塞进一个 monolithic prompt,导致行为脆弱、难调试、失败就要整段重跑。论文提出 runtime-structured task decomposition——任务切分与执行流改由可执行控制逻辑管理,LLM 只承担聚焦的判断类子任务,输出先过 schema 校验再进入下游。两组 SE 工作负载(Kubernetes 根因分析、多文件调试)× 三种配置(单体 / 静态分解 / 运行时结构化分解)× 每配置 10 次运行的对比显示:分解本身并不必然省钱——静态分解在 RCA 上重试成本 1,632±145 tokens,反而高于单体的 904±17,因为失败会连带重跑下游子任务;运行时结构化方案只重跑失败子任务,把重试成本压到 436±132 / 460 tokens,较单体最高降 51.7%,较静态分解降 73.2%。对 agent harness 设计者的直接启示:可靠性来自控制流结构,而非更长的 prompt。
为什么值得关注
用受控实验(2 工作负载 × 3 配置 × 10 runs)量化了一个 harness 设计常见误区:静态分解的重试成本会超过单体;只重跑失败子任务的运行时结构化分解降 51.7%-73.2%。做 coding agent 架构选型时可直接引用的数字。
要点摘录:
- 来源:arXiv 论文页面元数据 + 摘要(opencli 拉取)
- 标签:llm-agent, coding-agent, field-note
- 日期:2026-05-14
- 发表:ACM Conference on AI and Agentic Systems 2026 · Agentic Software Engineering workshop
- arXiv 分类:cs.SE (primary), cs.AI
关键信息
- 标题:Runtime-Structured Task Decomposition for Agentic Coding Systems
- URL:https://arxiv.org/abs/2605.15425
- 抓取日期:2026-09-10
English Abstract / Excerpt
Agentic coding systems increasingly use large language models (LLMs) for software engineering tasks such as debugging, root cause analysis, and code review. However, many existing systems encode task logic, execution flow, and output generation inside monolithic prompts. [...] We present runtime-structured task decomposition, an architectural approach in which task partitioning and execution flow are managed through executable control logic rather than prompt structure alone. LLMs are used only for focused judgment tasks, and outputs are validated against predefined schemas before downstream execution. [...] In the Kubernetes root cause analysis workload, the static decomposition baseline produced a retry cost of 1,632 +/- 145 tokens versus 904 +/- 17 tokens for the monolithic baseline because failures forced reruns of downstream subtasks. [...] The runtime-structured approach reran only failed subtasks, reducing retry costs to 436 +/- 132 tokens for root cause analysis and 460 tokens for debugging. Overall, the approach achieved up to 51.7% lower retry cost than monolithic systems and 73.2% lower retry cost than static decomposition baselines.
Obsidian Notes
- 内容由
opencli arxiv paper 2605.15425拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。