Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- ID: 6862e0fb
- 原文链接: https://arxiv.org/abs/2608.31046
- PDF: https://arxiv.org/pdf/2608.31046v1
- 作者: Yi Ding, Ruqi Zhang
- 日期: 2026-08-31
- 抓取时间: 2026-09-02T15:31:00Z
- 分类: learning
- 来源类型: paper
- 语言: zh
- 标签: distillation, post-training, rlvr, muon, qwen
- 质量评分: 4/5
中文导读
On-policy distillation(OPD)给 student 自采样轨迹做 token 级打分,看起来比 RLVR 的稀疏奖励更稠密,但作者量化分析发现 teacher 监督里噪声很多且随 teacher 规模增大而增多;意外的是 student 对这些噪声完全无感,留着或去掉噪声监督收敛差不多继续拆解下去,学习集中在低 log-probability token,用一个固定的负 advantage 就能匹配 teacher 给出的性能于是作者提出完全不需要 teacher 的 OPSA:熵自适应负 advantage 抑制尾部 token把概率质量在头部重新分配对 Qwen3-1.7B 基座,AIME24 的 Avg@32 提升 35.41 分(相对 +263%),比 OPD 还高 16.77 分,Pass@32 在三个 benchmark 全部翻倍这篇等于把蒸馏收益归因到"压制低概率 token"上,工程上意味着现有蒸馏 pipeline 可能一直在为不需要的 teacher 成本付费
一句话点评
On-policy distillation(OPD)给 student 自采样轨迹做 token 级打分,看起来比 RLVR 的稀疏奖励更稠密,但作者量化分析发现 teacher 监督里噪声很多且随 teacher 规模增大而增多.
English Abstract / Summary
The authors quantify teacher supervision in OPD and find substantial noise whose prevalence grows with teacher scale; surprisingly the student is insensitive to such noise, with convergence essentially unchanged whether noisy supervision is kept or removed. They trace gains to concentration of learning on low log-probability tokens and propose OPSA, a teacher-free method using entropy-adaptive negative advantage to suppress tail tokens. On Qwen3-1.7B base, OPSA lifts AIME24 Avg@32 by 35.41 points (+263% relative), beating OPD by 16.77 points and more than doubling Pass@32 on three benchmarks.
Obsidian Notes
- 由
daily-intake-evening2026-09-02 cron 从当日 Obsidian 摘要(论文流水线 / AK-RSS / ClawFeed / X 书签消化)发现并入库存量阶段。 - 中文导读与判断均锚定在条目已有摘要、源页面正文、作者、日期与分类信息;未补充源页面之外的实验细节。