TTPO: Test-Time Policy Optimization
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- RL 与 OPSD 等后训练方法依赖真值标签,因而无法做测试时训练(TTT);用多数投票伪标签替代又很脆弱:一次错误投票会污染教师并误导所有 token本文观察到失败模式的不对称性:与伪标签不一致的 rollout 无论投票对错几乎都是错的据此提出 TTPO 非对称目标:用 OPSD 蒸馏与伪标签一致的 rollout,用分组 RL 惩罚不一致的 rollout;token 级选择进一步细化两支:蒸馏下调已收敛位置的权重,RL 只惩罚自信的错误即使伪标签频繁出错,两支更新仍有依据;随着模型变强,多数投票路由带来更紧的自监督无标签条件下 TTPO 在五个竞赛级基准上与有标签 OPSD 持平,Qwen3-1.7B 在 TTT 设置下从 38.0% 提升到 45.2%,no-thinking 设置提升 +25.2% 到 +36.4%,并展现跨任务泛化
论文信息
- arXiv ID: 2608.27448
- 作者: Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu 等(11 人)
- 发表: 2026-08-27(更新:2026-08-27)
- 分类: cs.CL
- 备注: Project Page: https://zju-real.github.io/TTPO Code: https://github.com/ZJU-REAL/TTPO
- 原文链接: https://arxiv.org/abs/2608.27448
- PDF: https://arxiv.org/pdf/2608.27448v1
- 标签:
test-time-trainingpseudo-labelsopsdreinforcement-learningmath-reasoning
Abstract(原文)
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.
核心要点(英文摘要的中文提炼)
- 论文题为 TTPO: Test-Time Policy Optimization,发表于 arXiv(cs.CL,2026-08-27 提交,2026-08-27 更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
- 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。