Rethinking On-Policy Distillation of Large Language Models II: One Training Example
Original: https://arxiv.org/abs/2609.04172
Source platform: arxiv
Author: Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo...
Original date: 2026-09-03
summary_zh
Zixuan Fu 等 12 人(浙大 + 字节 + DeepSeek,2026-09-03 提交 cs.AI,29 页 20 图)把 OPD 推到数据极小极限单条 query 训练,单次 OPD 在数百步内仍能持续提升,恢复出接近 full-data OPD 跨任务域和模型家族的大部分增益解释方式是 state coverage:单条 query 的 rollout 100 步内覆盖 71.5% 状态;16 条语义不同的 query 覆盖到 98.9%,匹配 full-data结论:OPD data-overfed 但 algorithm-starved
summary_en
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data traini
one_liner
和 Sequential Beats Joint 同周同方向的姐妹篇:先 OPD 用 16 query 起步 + RL 锐化,比堆 SFT 数据再上 RL 更划算