Agent 与自动化 4.0 · 优秀 2026-08-25 · 论文

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

递归自我改进(RSI)在长程任务中一直很难:不断增长的历史会遮蔽任务状态并使技能调用失准论文提出 Recuris,一种面向长程 agent harness 的递归式经验-工作记忆架构:工作记忆跟踪任务进度并据此从经验记忆中引导技能选择,使技能使用锚定当前需求而非完整历史;这种耦合还把执行过程转化为结构化证据,可将失败定位到具体记忆组件一个固定的 Meta-Agent 把这些证据转化为局部化的经验证门控的技能记忆更新,更新重塑执行又产生新证据,形成有边界的递归记忆演化循环跨四个长程基准十个模型的评测中,37 个已完成的模型-基准配对里 35 个任务成功率获得提升

打开原文回到归档

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

  • ID: 7192217a
  • 原文链接: https://arxiv.org/abs/2608.24876
  • PDF: https://arxiv.org/pdf/2608.24876v1
  • 作者: Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
  • 日期: 2026-08-25
  • 更新: 2026-08-25
  • 分类: agents
  • 来源类型: paper
  • 标签: agents, memory, recursive-self-improvement, long-horizon, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-27T05:21:12Z

中文导读

递归自我改进(RSI)在长程任务中一直很难:不断增长的历史会遮蔽任务状态并使技能调用失准论文提出 Recuris,一种面向长程 agent harness 的递归式经验-工作记忆架构:工作记忆跟踪任务进度并据此从经验记忆中引导技能选择,使技能使用锚定当前需求而非完整历史;这种耦合还把执行过程转化为结构化证据,可将失败定位到具体记忆组件一个固定的 Meta-Agent 把这些证据转化为局部化的经验证门控的技能记忆更新,更新重塑执行又产生新证据,形成有边界的递归记忆演化循环跨四个长程基准十个模型的评测中,37 个已完成的模型-基准配对里 35 个任务成功率获得提升

为什么值得关注

Recuris:工作记忆与经验记忆耦合形成有边界的递归记忆演化循环,10 模型 x 4 长程基准中 35/37 配对成功率提升

关键信息

  • 论文标题: Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
  • 作者: Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
  • arXiv: https://arxiv.org/abs/2608.24876
  • 发布时间: 2026-08-25
  • arXiv 分类: cs.AI, cs.CL
  • 备注: Code: https://github.com/Gen-Verse/Recuris
  • 关联标签: agents, memory, recursive-self-improvement, long-horizon, arxiv
  • Code: https://github.com/Gen-Verse/Recuris
  • Evaluated on 4 long-horizon benchmarks x 10 models; tau-bench GPT-5.6 Sol +17.8 pts, Claude Opus 5 +15.6 pts to 87.9%

English Abstract

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris

Obsidian Notes

  • Metadata and abstract fetched via opencli arxiv paper 2608.24876 -f json (2026-08-27T05:21:12Z); response parsed list-or-dict tolerant.
  • 中文导读与价值判断锚定在条目已有摘要与论文摘要上,未补充摘要之外的实验细节。
  • Canonical backfill page for existing entry 7192217a in data/entries.json; generated by AAIF content-fetcher (2026-08-27), no push.