Agent 与自动化 4.0 · 优秀 2026-09-17 · 论文

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

agentic RL 训练方法:多轮 agent 只有 trajectory 级标量奖励,用自教师做 on-policy distillation 补 token 级监督;发现特权信息不保证教师可靠且教师收益随阶段变化RetireOPD 用 Adaptive Retirement 让学生在差距停止缩小并达到教师成功率目标比例后自行退休教师Qwen2.5 1.5B-7B 上 ALFWorld 成功率较 RL 基线提升 14.1%-18.8%

打开原文回到归档

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

  • ID: 91d86bee
  • 原文链接: https://arxiv.org/abs/2609.20784
  • PDF: https://arxiv.org/pdf/2609.20784v1
  • 作者: Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen
  • 日期: 2026-09-17
  • 更新: 2026-09-17
  • 分类: agents
  • 来源类型: paper
  • 标签: agentic-rl, distillation, multi-turn-agents, training
  • 质量评分: 4/5
  • 抓取时间: 2026-09-20T04:24:27Z

中文导读

agentic RL 训练方法:多轮 agent 只有 trajectory 级标量奖励,用自教师做 on-policy distillation 补 token 级监督;发现特权信息不保证教师可靠且教师收益随阶段变化RetireOPD 用 Adaptive Retirement 让学生在差距停止缩小并达到教师成功率目标比例后自行退休教师Qwen2.5 1.5B-7B 上 ALFWorld 成功率较 RL 基线提升 14.1%-18.8%

为什么值得关注

agentic RL 训练方法:多轮 agent 只有 trajectory 级标量奖励,用自教师做 on-policy distillation 补 token 级监督.

关键信息

  • 论文标题:RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
  • 作者:Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen
  • arXiv:https://arxiv.org/abs/2609.20784
  • 发布时间:2026-09-17
  • arXiv 分类:cs.CL, cs.AI
  • 关联标签:agentic-rl, distillation, multi-turn-agents, training
  • 备注:N/A

English Abstract

Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.

English Summary

Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读、价值判断、关键事实均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 抓取时间:2026-09-20T04:24:27Z
  • 抓取来源:opencli arxiv paper 2609.20784