模型与实验室 4.0 · 优秀 2026-08-27 · 论文

Boosting LLM Exploration via Weak-Model Guidance in RLVR

RLVR 训练常伴随策略熵下降,推理覆盖收窄大 k 的 pass@k 退化本文不走算法正则化路线,而是引入跨模型非参数扰动:强迫目标模型基于更小更弱模型生成的部分推理轨迹续写,陌生前缀打散过度自信鼓励探索不同推理路径数学基准上全面超过 vanilla RLVR,且 k 越大增益越明显,说明推理覆盖被实质拓宽;无需额外 SFT复杂奖励设计或提示工程即可缓解熵坍缩

打开原文回到归档

Boosting LLM Exploration via Weak-Model Guidance in RLVR

External-scan entry · 20260830 · awesome-ai-field-notes

中文摘要

RLVR 训练常伴随策略熵下降,推理覆盖收窄、大 k 的 pass@k 退化。本文不走算法正则化路线,而是引入跨模型非参数扰动:强迫目标模型基于更小、更弱模型生成的部分推理轨迹续写,陌生前缀打散过度自信、鼓励探索不同推理路径。数学基准上全面超过 vanilla RLVR,且 k 越大增益越明显,说明推理覆盖被实质拓宽;无需额外 SFT、复杂奖励设计或提示工程即可缓解熵坍缩。

English Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet effective approach to preserve the generative diversity of LLMs during RLVR. Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning trajectories generated by a smaller, weaker language models. These unfamiliar prefixes effectively disrupt over-confidence and encourage the exploration of distinct reasoning paths. We empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training. Experiments across multiple mathematical benchmarks show that our method consistently outperforms vanilla RLVR. Notably, the performance gain becomes increasingly pronounced as $k$ scales up, demonstrating a substantial expansion of reasoning coverage. Furthermore, our approach efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。