Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
原文链接: https://arxiv.org/abs/2609.31619
作者: Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata...
发布时间: 2026-09-25
源: arXiv外部扫描 (2026-09-29)
摘要
arXiv 2609.31619 发现:不用任何长度惩罚或早停机制,只让推理模型在自身思维链的中间点自监督地预测答案置信度(仅 600 个训练题),就能在 GemmaQwenNemotronGPT-OSS 四个模型族上把数学/科学/代码推理的生成 token 最多减少 25%,精度持平置信度只作为训练目标出现,损失函数里没有任何关于长度或停止的项;推理时用的也是标准生成流程对推理过程的分析显示置信度监督大体保留了基模型的高层推理构成,而不是选择性压制某些行为结论:高效推理可以作为学习元认知信号的副产品自然涌现,无需直接优化
English Summary
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: confidence. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping....
为什么值得关注
扩展 AAIF 对应主题线