Agent 与自动化 4.0 · 优秀 2026-08-31 · 论文

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

论文把长链路 agent RL 里的监督信用鸿沟讲清楚:PI(训练期可见的特权信息)能改写策略偏好,但 PI 引起的似然偏移并不直接回答一个可执行动作应分到多少 outcome 信用TASPO 从成功经验里构造对当前决策真正可用的 PI,在可执行动作粒度上聚合 PI 引起的偏好变化,再把相对动作支持转成正有界保持均值的轨迹优势权重在 3 个 agentic benchmark 上 TASPO 比 GRPO 提升 10.6%,并对未见任务有更好的泛化

打开原文回到归档

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

  • ID: cb6f8529
  • 原文链接: https://arxiv.org/abs/2608.31077
  • PDF: https://arxiv.org/pdf/2608.31077v1
  • 作者: Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, Xuhong Zhang
  • 日期: 2026-08-31
  • 更新: 2026-08-31
  • 分类: agents
  • 来源类型: paper
  • 标签: llm-agent, rl, credit-assignment, agentic-rl
  • 质量评分: 4/5
  • 抓取时间: 2026-09-02T12:28:56+00:00

中文导读

论文把长链路 agent RL 里的监督信用鸿沟讲清楚:PI(训练期可见的特权信息)能改写策略偏好,但 PI 引起的似然偏移并不直接回答一个可执行动作应分到多少 outcome 信用TASPO 从成功经验里构造对当前决策真正可用的 PI,在可执行动作粒度上聚合 PI 引起的偏好变化,再把相对动作支持转成正有界保持均值的轨迹优势权重在 3 个 agentic benchmark 上 TASPO 比 GRPO 提升 10.6%,并对未见任务有更好的泛化

为什么值得关注

论文把长链路 agent RL 里的 supervision-credit gap 拆开:PI(训练期可见的特权信息)能改写策略偏好,但 PI 引起的代变偏移并不直接回答一个可执行动作应分到多少 outcome 信用。TASPO 从成功经验里构造 decision-applicable PI,在可执行动作粒度上聚合 PI 引起的偏好变化,再把相对动作支持转成正有界、保持均值的轨迹优势权重,让 verified outcome 决定更新方向与平均幅度,PI 只负责在动作间再分配信用。在 3 个 agentic benchmark 上 TASPO 比 GRPO 提升 10.6%,并在未见任务上表现出更好的泛化。论文标注 Work in progress,后续需关注实验细节与更大规模基准的复现。

关键信息

  • 论文标题: Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization
  • 作者: Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, Xuhong Zhang
  • arXiv: https://arxiv.org/abs/2608.31077
  • 发布时间: 2026-08-31
  • 最后更新: 2026-08-31
  • arXiv 分类: cs.AI (cs.AI)
  • 论文备注: Work in progress
  • 关联标签: llm-agent, rl, credit-assignment, agentic-rl

English Abstract

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at a token granularity misaligned with executable decisions, and lack the outcome semantics required for reinforcement. We introduce TASPO, which converts privileged supervision into outcome-grounded action credit. TASPO constructs decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. Thus, the verified outcome determines the update direction and average scale, while PI only redistributes credit across actions. Across three agentic benchmarks, TASPO improves over GRPO by 10.6\% and generalizes better to unseen tasks. Further analysis indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process. These findings offer the community another interesting perspective.

English Summary

The paper frames the supervision-credit gap in outcome-based agent RL: privileged information (PI) re-evaluates sampled behavior at training time, but PI-induced likelihood shifts only say how extra information changes policy preference, not how an executable action should inherit the verified outcome. TASPO builds decision-applicable PI from successful experience, aggregates PI-induced shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the trajectory advantage. Reported gain over GRPO: +10.6% across three agentic benchmarks, with better generalization to unseen tasks.

Obsidian Notes

  • 内容由 opencli arxiv paper 2608.31077 -f json 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。