Agent 与自动化 5.0 · 必读 2026-08-24 · 论文

Prime Agent: A Self-Improving RLM Harness

Prime Agent 是开源的长视野评测与 coding-agent harness:持久 IPython REPL 承担 Recursive Language Model 抽象,负责程序化的上下文处理与 test-time compute;Continual Harness 让历史记忆技能prompt子 agent 规格跨轨迹保留...

打开原文回到归档

Prime Agent: A Self-Improving RLM Harness

  • ID: 25fd4d45
  • 原文链接: https://arxiv.org/abs/2608.23552
  • PDF: https://arxiv.org/pdf/2608.23552v1
  • 作者: S, e, t, h, , K, a, r, t, e, n, ,, , A, l, e, x, , L, ., , Z, h, a, n, g, ,, , K, e, v, i, n, , T, h, o, m, a, s, ,, , S, e, b, a, s, t, i, a, n, , M, ü, l, l, e, r, ,, , E, l, i, e, , B, a, k, o, u, c, h, ,, , D, a, n, i, e, l, , A, u, r, a, s, ,, , M, i, k, a, , S, e, n, g, h, a, a, s, ,, , F, a, r, e, s, , O, b, e, i, d, ,, , K, o, n, s, t, a, n, t, i, n, , D, u, n, a, s, ,, , J, o, h, a, n, n, e, s, , H, a, g, e, m, a, n, n, ,, , S, a, m, i, , J, a, g, h, o, u, a, r
  • 日期: 2026-08-24
  • 更新: 2026-08-24
  • 分类: agents
  • 来源类型: paper
  • 标签: prime-agent, rlm, agent-harness, arc-agi, test-time-compute
  • 质量评分: 5/5
  • 抓取时间: 2026-08-26T04:27:05Z

中文导读

Prime Agent 是开源的长视野评测与 coding-agent harness:持久 IPython REPL 承担 Recursive Language Model 抽象,负责程序化的上下文处理与 test-time compute;Continual Harness 让历史记忆技能prompt子 agent 规格跨轨迹保留;递归子 agent 之间直接通信,Agents View 供人检查和接管 daemon 化的会话harness 统一处理执行恢复验证与资源记账,把策略构造留给模型本身,避免 harness 失败被记成模型失败效果上,ARC-AGI-3 RHAE Best@1 从 30% 提到 95.5%,在长上下文编码GPU kernel 生成模拟器构建自主 nanoGPT speedrun 等任务上追平或超过原生与主流 harness;Factorio 上精炼循环支撑连续科技演进专职子 agent 支持并行分工代码开源于 github.com/PrimeIntellect-ai/prime-agent

为什么值得关注

开源 RLM harness:持久 IPython REPL + 跨轨迹记忆/技能,把 ARC-AGI-3 RHAE Best@1 从 30% 拉到 95.5%

对做 agent harness 与长视野评测的工程师,这是一个可直接上手的开源参考实现:持久 IPython REPL、跨轨迹保留的记忆与技能、人可检查接管的 Agents View,代码已在 GitHub 开源(PrimeIntellect-ai/prime-agent)。

关键信息

  • 论文标题: Prime Agent: A Self-Improving RLM Harness
  • 作者: S, e, t, h, , K, a, r, t, e, n, ,, , A, l, e, x, , L, ., , Z, h, a, n, g, ,, , K, e, v, i, n, , T, h, o, m, a, s, ,, , S, e, b, a, s, t, i, a, n, , M, ü, l, l, e, r, ,, , E, l, i, e, , B, a, k, o, u, c, h, ,, , D, a, n, i, e, l, , A, u, r, a, s, ,, , M, i, k, a, , S, e, n, g, h, a, a, s, ,, , F, a, r, e, s, , O, b, e, i, d, ,, , K, o, n, s, t, a, n, t, i, n, , D, u, n, a, s, ,, , J, o, h, a, n, n, e, s, , H, a, g, e, m, a, n, n, ,, , S, a, m, i, , J, a, g, h, o, u, a, r
  • arXiv: https://arxiv.org/abs/2608.23552
  • 发布时间: 2026-08-24
  • arXiv 分类: c, s, ., A, I, ,, , c, s, ., C, L, ,, , c, s, ., S, E
  • 备注: 16 pages, 10 figures. Technical report. Code: https://github.com/PrimeIntellect-ai/prime-agent
  • 关联标签: prime-agent, rlm, agent-harness, arc-agi, test-time-compute

English Abstract

Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns. On Factorio, we find refinement allows for continuous technology progression and dedicated subagents enable parallelized work. Code is available at https://github.com/PrimeIntellect-ai/prime-agent.

English Summary

Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute; Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories; recursive subagents coordinate through direct agent-to-agent communication; an Agents View lets humans inspect and manage daemon-backed sessions. The harness standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model, so harness failures do not masquerade as model failures....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页由 AAIF content-fetcher 定时任务生成(2026-08-26),仅新增内容页,未改动 entries.json。