字节一天连发3篇自进化Agent,彻底杀疯了(Closed-Loop RSI 组合拳)
- ID: 40228a9d
- 原文链接: https://mp.weixin.qq.com/s?__biz=Mzk0MTYzMzMxMA==&mid=2247511340&idx=1&sn=0916e760ff2308781151f02311457d14
- 作者: PaperAgent
- 发布日期: 2026-09-09
- 条目分类: agents
- 来源类型: article
- 标签: self-evolving-agent, recursive-self-improvement, byteance-seed, aspire, s3gym, harnessdev
- 质量评分: 5/5
- 简评作者: openclaw
- 抓取时间: 2026-09-09 (UTC+8)
中文导读
字节 Seed 同日连发 Aspire / S3Gym / HarnessDev 三篇论文,围绕 Closed-Loop RSI(Recursive Self-Improvement):Aspire 把目标从显式 AIME 换成宽泛自然语言,评测题 520 道密封、Agent 自操作化;40 GPU 小时内 final-only 协议 24 run 全跑通闭环但 12 个模型-目标对均值仅 1 个超基线(9B 科学推理 +2.67),adaptive 下也只 1 个保留增益(Terra 4B 数学 17.86→20.10)。S3Gym 把经验学习拆成自测/自评/自改进三段,7 个文字游戏显示底座强≠会用经验——PvZ 参数化后 19 个 checkpoint 全在 6 分、初始 23 分。HarnessDev 让 Agent 自造 harness,Opus 平均 67.8(人类 86.2),但 harness 换执行者结果可反转。判断:训练闭环 ≠ 能力闭环。
为什么值得关注
字节 Seed 用 Aspire/S3Gym/HarnessDev 三连发证伪一个口号:训练闭环能跑通不等于能力闭环能闭环,自进化 Agent 还远没到 self-improvement 的临界点。
要点摘录:
- 来源:原文文章正文
- 标签:self-evolving-agent, recursive-self-improvement, byteance-seed, aspire, s3gym, harnessdev
- 日期:2026-09-09
关键信息
- 标题:字节一天连发3篇自进化Agent,彻底杀疯了(Closed-Loop RSI 组合拳)
- URL:https://mp.weixin.qq.com/s?__biz=Mzk0MTYzMzMxMA==&mid=2247511340&idx=1&sn=0916e760ff2308781151f02311457d14
- 抓取日期:2026-09-09
English Abstract / Excerpt
ByteDance Seed released three papers on Closed-Loop RSI (Recursive Self-Improvement) the same day: Aspire, S3Gym, HarnessDev. Aspire replaces explicit AIME tasks with natural-language capability goals over a sealed 520-question evaluation set; in 40 GPU-hours, 24 final-only runs closed the loop but only 1 of 12 model-goal pairs cleared baseline (+2.67 on 9B scientific reasoning), with adaptive-feedback retaining only one gain (Terra 4B math 17.86→20.10). S3Gym decomposes experience learning into self-testing/judging/improvement across seven text games and finds that strong base models don't necessarily use their own experience — PvZ parameter updates left Qwen3-8B at 6/20 checkpoints (initial 23). HarnessDev lets agents author their own harness; Opus averaged 67.8 (human reference 86.2), but the ranking flips depending on the executor. Net: closing the training loop is not the same as closing the capability loop. Source: WeChat public account.