模型与实验室 4.0 · 优秀 2026-09-28 · 论文

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

提出 Retrospection-Only Fine-Tuning(ROFT):智能体完成任务后,仅对自己的复盘解释做 next-token 微调,不用外部教师不用奖励策略更新在 Qwen3.5-4B 的软件工程实验中,用混合成败的基础模型轨迹训练 20 步后,SWE-bench Verified 达 49.2%Pro 达 26.8%(同评测设置下 GRPO 40 步为 48.0%/25.3%),且训练前期进度更快它还能解出基础模型 64 次采样全部失败的任务,说明无需任何成功轨迹即可开始学习行为分析显示复盘间接起到信用分配作用;提示复盘偏向更直接的解法还能缩短后续尝试结论:学会解释本身就能迁移为学会行动

打开原文回到归档

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

  • ID: 6352ce00
  • 原文链接: https://arxiv.org/abs/2609.35741
  • PDF: https://arxiv.org/pdf/2609.35741v1
  • 作者: Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim
  • 日期: 2026-09-28
  • 更新: 2026-09-28
  • 分类: models
  • 来源类型: paper
  • 标签: $agentic-rl, $fine-tuning, $swe-bench, $self-improvement
  • 质量评分: 4/5
  • 抓取时间: 2026-09-30T04:26:15Z

中文导读

提出 Retrospection-Only Fine-Tuning(ROFT):智能体完成任务后,仅对自己的复盘解释做 next-token 微调,不用外部教师不用奖励策略更新在 Qwen3.5-4B 的软件工程实验中,用混合成败的基础模型轨迹训练 20 步后,SWE-bench Verified 达 49.2%Pro 达 26.8%(同评测设置下 GRPO 40 步为 48.0%/25.3%),且训练前期进度更快它还能解出基础模型 64 次采样全部失败的任务,说明无需任何成功轨迹即可开始学习行为分析显示复盘间接起到信用分配作用;提示复盘偏向更直接的解法还能缩短后续尝试结论:学会解释本身就能迁移为学会行动

为什么值得关注

把「复盘解释」本身当作训练信号,在没有奖励模型和外部教师的情况下逼近 GRPO 的评测效果,为小模型智能体的自改进提供了极低成本的路线。

关键信息

  • 论文标题: Shockingly Simple Self-retrospection Improves Agentic Models Without RL
  • 作者: Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim
  • arXiv: https://arxiv.org/abs/2609.35741
  • 发布时间: 2026-09-28
  • arXiv 分类: cs.AI, cs.CL
  • Comments: 62 pages, 18 figures, 5 tables, including appendices
  • 关联标签: agentic-rl, fine-tuning, swe-bench, self-improvement

English Abstract

People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

English Summary

People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。