SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- ID: 250ac425
- 原文链接: https://arxiv.org/abs/2608.24870
- PDF: https://arxiv.org/pdf/2608.24870v1
- 作者: Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang
- 日期: 2026-08-25
- 更新: 2026-08-25
- 分类: learning
- 来源类型: paper
- 标签: rl, agentic-rl, policy-optimization, asynchronous-training, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-27T05:21:12Z
中文导读
组相对强化学习需要等待同一 prompt 的兄弟 rollout,这对长且长度差异大的工具调用轨迹来说成本很高此前的单流策略优化(SPO)用持续的 prompt 级价值估计摆脱了这一依赖,但其配方先把每条轨迹的优势做白化,再去优化 token 平均的 actor loss作者证明轨迹级居中通常并不能使 actor 实际消费的 token 加权量居中,并改用动作 token 度量下的标准化终局结果优势来修复这个错配;另外按生成证据的策略事件而非学习器接收顺序来组织 prompt 证据在与 ALFWorld 两个模型规模及 Math-TIR 的对齐对照运行中,SPO++ 相比 SPO 提升在线学习效率;配对消融显示动作-token-度量归一化是所测组件中最强的一项
为什么值得关注
SPO++:发现单流策略优化中轨迹白化与 token-mean actor loss 的度量错配,用动作 token 度量标准化修复,异步 agentic RL 效率提升
关键信息
- 论文标题: SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- 作者: Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang
- arXiv: https://arxiv.org/abs/2608.24870
- 发布时间: 2026-08-25
- arXiv 分类: cs.AI
- 备注: 9 pages, 2 figures, 3 tables
- 关联标签: rl, agentic-rl, policy-optimization, asynchronous-training, arxiv
- 9 pages, 2 figures, 3 tables
- Matched-run gains on ALFWorld (two model scales) and Math-TIR; paired ablation pins action-token-measure normalization as the strongest component
English Abstract
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. We show that trajectory centering generally does not center the token-weighted quantity consumed by the actor, and fix the mismatch by standardizing terminal-outcome advantages under the action-token measure. We additionally organize prompt evidence by the policy event that generated it rather than learner receipt order. Across matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO. A paired ablation identifies action-token-measure normalization as the strongest tested component.
Obsidian Notes
- Metadata and abstract fetched via
opencli arxiv paper 2608.24870 -f json(2026-08-27T05:21:12Z); response parsed list-or-dict tolerant. - 中文导读与价值判断锚定在条目已有摘要与论文摘要上,未补充摘要之外的实验细节。
- Canonical backfill page for existing entry
250ac425indata/entries.json; generated by AAIF content-fetcher (2026-08-27), no push.