Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
- ID: c9bc0041
- 原文链接: https://arxiv.org/abs/2609.38147
- PDF: https://arxiv.org/pdf/2609.38147v1
- 作者: Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal
- 日期: 2026-09-29
- 更新: 2026-09-29
- 分类: agents
- 来源类型: paper
- 标签: arxiv, inference-harness, meta-reasoning, controller-worker, test-time-compute
- 质量评分: 4/5
- 抓取时间: 2026-10-01T12:21:44+00:00
中文导读
随着 agent 承担更长、更复杂的问题,控制执行本身成为一项独立任务:运行中的每一步都带来新的控制选择——基于哪些已有部分继续、是否推翻重来、何时停止。论文提出 agentic meta-reasoning:一个推理时 harness,把这些运行控制决策变成显式、结构化的推理过程。Worker 承担任务级计算;controller 负责汇总本次运行已确立的进展、探索下一步选项、在剩余预算下估算每个选项的价值,并结合持久记忆中的上下文派发工作。决策之间,controller 只携带一份紧凑的运行账户(compact account),而不是重放全部历史。
基线覆盖生产级 coding agent、研究 harness,以及使用相同 worker 与相同算力预算的 Direct Control Agent 对照组。在通过程序重建测试长程 agentic 能力的 ProgramBench 上:GPT-5.5 搭配 meta-reasoning 达 71.5%(Codex 为 58.0%);Opus 4.8 达 67.2%(Claude Code 为 65.5%)。在抽象推理、多域长程推理、证明生成等其他基准上,对三个前沿模型平均提升 3.6–4.2 个百分点。在直接控制已经趋于平台的预算区间内 meta-reasoning 仍持续提升,但预算较小时其开销反而有害。Artifact-graph 分析显示:对早期工作的复用更多、多数设置下正确解覆盖更高、最终选择环节收益不均匀。结论:随着 agent 运行时长增长,把算力花在结构化控制上变得越来越重要。
为什么值得关注
controller-worker 形态的 agentic meta-reasoning:把运行控制决策变成显式推理,决策间只携带紧凑进度账户。它回答了 test-time compute 的下一个问题——算力不再只堆给"想得更久",而是堆给"管得更聪明";对 agent harness 设计者,这是把调度/重启/停止从隐式启发式升级为显式推理的完整配方与量化证据。
关键信息
- 论文标题:Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
- 作者:Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal
- arXiv:https://arxiv.org/abs/2609.38147
- 发布时间:2026-09-29
- arXiv 分类:cs.AI
- 关联标签:arxiv, inference-harness, meta-reasoning, controller-worker, test-time-compute
English Abstract
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.
English Summary
Introduces agentic meta-reasoning, an inference-time harness that makes run-control choices (which partial work to build on, whether to restart, when to stop) an explicit structured reasoning process. Workers perform task-level computation while a controller consolidates progress, explores options, values each under the remaining budget, and dispatches work with context from persistent memory, carrying only a compact account of the run between decisions. Baselines span production coding agents and research harnesses plus a matched Direct Control Agent.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。