Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- 进化策略(ES)作为省内存的 LLM 推理后训练范式,其优化行为此前研究不足本文系统分析 ES 动态与机制,理论和实证均证明 ES 带来更宽的推理覆盖:ES 种群内验证器投影的 Jensen-Shannon 多样性支撑更高 Pass@K;与出现熵坍缩的 GRPO 不同,ES 在提升 Pass@1 的同时取得比 GRPO 更高的 Pass@K作者进一步提出 GRPO-ES 顺序训练策略,结合 GRPO 的 Pass@1 优势与 ES 的 Pass@K 优势另一关键发现:尽管整模参数大幅漂移,ES 的任务增益只来自少数大幅更新的稀疏子集(功能稀疏性),held-out 评测未观察到灾难性遗忘超参方面,越大的 LLM 需要越小的种群规模结论:ES 是与 GRPO 互补的独立推理后训练范式,而非省内存的次优替代
论文信息
- arXiv ID: 2608.27351
- 作者: Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong 等(10 人)
- 发表: 2026-08-27(更新:2026-08-28)
- 分类: cs.LG
- 原文链接: https://arxiv.org/abs/2608.27351
- PDF: https://arxiv.org/pdf/2608.27351v2
- 标签:
evolution-strategiesgrpopass-at-kentropy-collapsepost-training
Abstract(原文)
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.
核心要点(英文摘要的中文提炼)
- 论文题为 Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO,发表于 arXiv(cs.LG,2026-08-27 提交,2026-08-28 更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
- 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。