模型与实验室 4.0 · 优秀 2026-08-27 · 论文

TTPO: Test-Time Policy Optimization

RL 与 OPSD 等后训练方法依赖真值标签,因而无法做测试时训练(TTT);用多数投票伪标签替代又很脆弱:一次错误投票会污染教师并误导所有 token本文观察到失败模式的不对称性:与伪标签不一致的 rollout 无论投票对错几乎都是错的据此提出 TTPO 非对称目标:用 OPSD 蒸馏与伪标签一致的 rollout,用分组 RL 惩罚不一致的 rollout;token 级选择进一步细化两支:蒸馏下调已收敛位置的权重,RL 只惩罚自信的错误即使伪标签频繁出错,两支更新仍有依据;随着模型变强,多数投票路由带来更紧的自监督无标签条件下 TTPO 在五个竞赛级基准上与有标签 OPSD 持平,Qwen3-1.7B 在 TTT 设置下从 38.0% 提升到 45.2%,no-thinking 设置提升 +25.2% 到 +36.4%,并展现跨任务泛化

打开原文回到归档

TTPO: Test-Time Policy Optimization

摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
  • RL 与 OPSD 等后训练方法依赖真值标签,因而无法做测试时训练(TTT);用多数投票伪标签替代又很脆弱:一次错误投票会污染教师并误导所有 token本文观察到失败模式的不对称性:与伪标签不一致的 rollout 无论投票对错几乎都是错的据此提出 TTPO 非对称目标:用 OPSD 蒸馏与伪标签一致的 rollout,用分组 RL 惩罚不一致的 rollout;token 级选择进一步细化两支:蒸馏下调已收敛位置的权重,RL 只惩罚自信的错误即使伪标签频繁出错,两支更新仍有依据;随着模型变强,多数投票路由带来更紧的自监督无标签条件下 TTPO 在五个竞赛级基准上与有标签 OPSD 持平,Qwen3-1.7B 在 TTT 设置下从 38.0% 提升到 45.2%,no-thinking 设置提升 +25.2% 到 +36.4%,并展现跨任务泛化

论文信息

Abstract(原文)

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.

核心要点(英文摘要的中文提炼)

  • 论文题为 TTPO: Test-Time Policy Optimization,发表于 arXiv(cs.CL,2026-08-27 提交,2026-08-27 更新)。
  • 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
  • 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。