模型与实验室 4.0 · 优秀 2026-08-19 · 论文

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

诊断出 on-policy 蒸馏(OPD)在长上下文证据聚合任务上的失配:token 级教师信号偏爱局部合理但遗漏分散证据违反全局约束的回答,与响应级 verifier 奖励的分歧随输入变长逐步拉大提出 GC-OPD:在每个 rollout 组内分别归一化 verifier 奖励与 OPD 轨迹分,把差值作为带符号的教师-verifier 分歧残差,再用相对优势做 token 级分摊五个长上下文基准:Qwen3-4B 五基准均分 29.0840.47Qwen3-8B 35.1244.65(vanilla OPD 为 39.31/43.56)无需额外模型只在组内归一化,工程上好落地

打开原文回到归档

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses from two representative long-context evidence-aggregation tasks. Across longer input ranges, trajectory-level OPD scores become progressively less aligned with verifier rewards, indicating teacher-verifier disagreement. Motivated by this observation, we introduce Group-Calibrated On-Policy Distillation (GC-OPD). GC-OPD separately normalizes verifier rewards and trajectory-level OPD scores within each rollout group and uses their difference as a signed teacher-verifier disagreement residual. Relative-advantage-based credit assignment (RACA) distributes this trajectory-level residual across tokens according to their relative OPD advantages while preserving the original OPD signal. Across five long-context benchmarks, post-training with GC-OPD raises the five-benchmark averages of the official Qwen3-4B and Qwen3-8B checkpoints from 29.08 to 40.47 and from 35.12 to 44.65, respectively. Vanilla OPD reaches 39.31 and 43.56 under the same setup. Controlled ablations show that the signed residual is more effective than either an additional OPD-derived term or direct group-normalized verifier reward addition, while RACA further improves over uniform token allocation. Together, these results demonstrate that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance. Code is available at https://github.com/SolereZhang/GC-OPD.

Authors: Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou Published: 2026-08-19 Categories: cs.LG, cs.AI, cs.CL arXiv: 2608.19181

Source: https://arxiv.org/abs/2608.19181
Captured: 2026-08-21 (AAIF daily-intake-evening)