Agent 与自动化 4.0 · 优秀 2026-09-21 · 论文

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

LLM agent 的能力很大程度由 harness(围绕冻结底座的提示词控制流工具记忆与上下文管理)放大;近年方法通过迭代提出并筛选 harness 的组件级修改,形成 agent 系统层的递归自我改进(RSI),但这种进化会过拟合训练任务分布内大涨分布外收益缩水甚至消失RRSI 把正则化原则引入 harness 自我进化:提议端用时间退火的修改预算限制单次打包的改动数,并基于进化历史鼓励未探索轨迹...

打开原文回到归档

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

  • ID: 1fd62c8a
  • 原文链接: https://arxiv.org/abs/2609.24972
  • PDF: https://arxiv.org/pdf/2609.24972v1
  • 作者: Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
  • 日期: 2026-09-21
  • 更新: 2026-09-21
  • 分类: agents
  • 来源类型: paper
  • 标签: arxiv, agent-harness, recursive-self-improvement, evaluation, google
  • 质量评分: 4/5
  • 抓取时间: 2026-09-23T04:27:36+00:00

中文导读

LLM agent 的能力很大程度由 harness(围绕冻结底座的提示词控制流工具记忆与上下文管理)放大;近年方法通过迭代提出并筛选 harness 的组件级修改,形成 agent 系统层的递归自我改进(RSI),但这种进化会过拟合训练任务分布内大涨分布外收益缩水甚至消失RRSI 把正则化原则引入 harness 自我进化:提议端用时间退火的修改预算限制单次打包的改动数,并基于进化历史鼓励未探索轨迹;选择端配 critic 过滤针对特定 benchmark 的提议pruner 剔除过小/过贵/失效的改动,偏向可复用的 agent 机制在编码agentic 工作区工程设计共 8 个 benchmark 上,进化目标集最高 +14.1 分,5 个 OOD benchmark 最高 +4.7 分,policy token 比无正则进化少 30%代码在 github.com/google-research/rrsi

为什么值得关注

RRSI:给 agent harness 递归自我进化加正则化,OOD benchmark 仍 +4.7 分,policy token 省 30%

关键信息

  • 论文标题:RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
  • 作者:Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
  • arXiv:https://arxiv.org/abs/2609.24972
  • 发布时间:2026-09-21
  • arXiv 分类:cs.LG, cs.AI, cs.CL
  • 评论/页数:N/A
  • 关联标签:arxiv, agent-harness, recursive-self-improvement, evaluation, google

English Abstract

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.

English Summary

An LLM agent's capability is magnified by its harness (prompts, control flow, tooling, memory, context management around a frozen backbone). Recent methods automate this by iteratively proposing and selecting component-wise harness edits, a form of recursive self-improvement (RSI) at the agent-system level, but such evolution can overfit: large in-distribution gains shrink or vanish on out-of-distribution benchmarks. RRSI adds regularization to harness self-improvement: the proposer operates with a temporally annealed edit budget and is encouraged toward unexplored trajectories; the selector gains a critic screening benchmark-specific proposals and a pruner dropping changes that are too small, too costly, or no longer useful, favoring reusable mechanisms....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。