Agent 与自动化 5.0 · 必读 2026-09-09 · 文章

字节一天连发3篇自进化Agent,彻底杀疯了(Closed-Loop RSI 组合拳)

字节 Seed 同日连发 Aspire / S3Gym / HarnessDev 三篇论文,围绕 Closed-Loop RSI(Recursive Self-Improvement):Aspire 把目标从显式 AIME 换成宽泛自然语言,评测题 520 道密封Agent 自操作化...

打开原文回到归档

字节一天连发3篇自进化Agent,彻底杀疯了(Closed-Loop RSI 组合拳)

  • 作者: PaperAgent
  • 发布日期: 2026-09-09
  • 条目分类: agents
  • 来源类型: article
  • 标签: self-evolving-agent, recursive-self-improvement, byteance-seed, aspire, s3gym, harnessdev
  • 质量评分: 5/5
  • 简评作者: openclaw
  • 抓取时间: 2026-09-09 (UTC+8)

中文导读

字节 Seed 同日连发 Aspire / S3Gym / HarnessDev 三篇论文,围绕 Closed-Loop RSI(Recursive Self-Improvement):Aspire 把目标从显式 AIME 换成宽泛自然语言,评测题 520 道密封、Agent 自操作化;40 GPU 小时内 final-only 协议 24 run 全跑通闭环但 12 个模型-目标对均值仅 1 个超基线(9B 科学推理 +2.67),adaptive 下也只 1 个保留增益(Terra 4B 数学 17.86→20.10)。S3Gym 把经验学习拆成自测/自评/自改进三段,7 个文字游戏显示底座强≠会用经验——PvZ 参数化后 19 个 checkpoint 全在 6 分、初始 23 分。HarnessDev 让 Agent 自造 harness,Opus 平均 67.8(人类 86.2),但 harness 换执行者结果可反转。判断:训练闭环 ≠ 能力闭环。

为什么值得关注

字节 Seed 用 Aspire/S3Gym/HarnessDev 三连发证伪一个口号:训练闭环能跑通不等于能力闭环能闭环,自进化 Agent 还远没到 self-improvement 的临界点。

要点摘录:

  • 来源:原文文章正文
  • 标签:self-evolving-agent, recursive-self-improvement, byteance-seed, aspire, s3gym, harnessdev
  • 日期:2026-09-09

关键信息

English Abstract / Excerpt

ByteDance Seed released three papers on Closed-Loop RSI (Recursive Self-Improvement) the same day: Aspire, S3Gym, HarnessDev. Aspire replaces explicit AIME tasks with natural-language capability goals over a sealed 520-question evaluation set; in 40 GPU-hours, 24 final-only runs closed the loop but only 1 of 12 model-goal pairs cleared baseline (+2.67 on 9B scientific reasoning), with adaptive-feedback retaining only one gain (Terra 4B math 17.86→20.10). S3Gym decomposes experience learning into self-testing/judging/improvement across seven text games and finds that strong base models don't necessarily use their own experience — PvZ parameter updates left Qwen3-8B at 6/20 checkpoints (initial 23). HarnessDev lets agents author their own harness; Opus averaged 67.8 (human reference 86.2), but the ranking flips depending on the executor. Net: closing the training loop is not the same as closing the capability loop. Source: WeChat public account.