Agent 与自动化 3.0 · 值得看 2026-09-29 · 论文

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific...

把 agent 规划失败拆成选错计划与执行走样两类,提出 Planning-as-Routing:模型声明四种规划模式之一(Predefined/Sequential/Hierarchical/Search),确定性路由器把任务派给对应模式的专用执行器四基准三模型的三条结论:通用 Plan+ReAct 在三个基准上只有 22-45% 的轨迹保住声明的规划结构长计划尤其差;模式有效性随环境与模型变化(ALFWorld 上 Search 最佳SWE-bench 上 Hierarchical 最佳同基准内最强模式可跨模型变化)...

打开原文回到归档

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

  • ID: 3c1d6de6
  • 原文链接: https://arxiv.org/abs/2609.38108
  • 作者: Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos López de Prado, Shadab Khan
  • 日期: 2026-09-29
  • 更新: N/A
  • 分类: agents
  • 来源类型: paper
  • 标签: agent-planning, execution, benchmark, react
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

把 agent 规划失败拆成选错计划与执行走样两类,提出 Planning-as-Routing:模型声明四种规划模式之一(Predefined/Sequential/Hierarchical/Search),确定性路由器把任务派给对应模式的专用执行器。四基准三模型的三条结论:通用 Plan+ReAct 在三个基准上只有 22-45% 的轨迹保住声明的规划结构、长计划尤其差;模式有效性随环境与模型变化(ALFWorld 上 Search 最佳、SWE-bench 上 Hierarchical 最佳、同基准内最强模式可跨模型变化);最大收益来自执行侧——专用执行器把 ALFWorld 成功率从 0.48 提到 0.92、SWE-bench Verified 从 0.36 提到 0.44。模型目前不能可靠挑出每题最强模式,few-shot 只在部分基准-模型组合里改善选择。

论文信息

  • arXiv ID: 2609.38108
  • 提交日期: 2026-09-29
  • arXiv 分类: cs.AI, cs.LG
  • 作者: Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos López de Prado, Shadab Khan
  • 链接: https://arxiv.org/abs/2609.38108

Abstract(arXiv 原文)

Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.

为什么值得关注

把「规划失败」拆成选择与执行两段是最大贡献——当下收益在执行侧、模式选择仍是开放问题,对 harness 设计者是直接输入。

English Summary

This paper decomposes agent planning failures into plan selection versus faithful execution and proposes Planning-as-Routing: the LLM declares one of four planning modes (Predefined/Sequential/Hierarchical/Search) and a deterministic router dispatches the task to pattern-specific executors. Across four benchmarks and three LLMs, generic Plan+ReAct preserves its declared structure in only 22-45% of trajectories (worse for long plans); executors lift ALFWorld success from 0.48 to 0.92 and SWE-bench Verified from 0.36 to 0.44, while reliable per-task mode selection by the model itself remains open.