Agent 与自动化 4.0 · 优秀 2026-09-03 · 论文

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

面向长程智能体训练的细粒度信用分配:RLVR 依赖可编程检查器,而多数长程任务没有;多准则 rubric 每条轨迹只打一个标量,跨几十步是贫信号DRACO 在训练中动态生成 rubric 以跟踪策略能力演进,每条完成的轨迹评分一次,再把该评判闭式重分配到负责相应 rubric 的步骤上,在 GRPO 中产生差异化的逐步优势,不引入任何需训练的归因模块;AppWorld 上超基线模型 15.9 分

打开原文回到归档

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

  • ID: 594a040e
  • 原文链接: https://arxiv.org/abs/2609.04094
  • PDF: https://arxiv.org/pdf/2609.04094v1
  • 作者: Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk
  • 日期: 2026-09-03
  • 更新: 2026-09-03
  • 分类: cs.AI, cs.LG, cs.SE
  • 来源类型: arxiv
  • 标签: rlvr, credit-assignment, rubrics, grpo, long-horizon, appworld, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-06T04:24:00Z

中文导读

RLVR 在任务自带可编程检查器时效果好,但多数长程智能体领域没有。在真值成功信号不可得的 outcome-blind 设定下,多准则 rubric 是常用的替代奖励,但每条轨迹只打一个标量,跨几十步是贫信号。DRACO 在训练过程中动态生成 rubric 以跟踪策略能力的演进,每条完成的轨迹评分一次,再把这一评判闭式重分配到负责相应 rubric 注释的步骤上,在 GRPO 中产生差异化的逐步优势,且不引入任何需要训练的归因模块。在 AppWorld 上比基座模型高 15.9 分、比使用稀疏真值奖励的 GRPO 高 5.3 分——而它自己没有用任何验证器;在域外 Tau-Bench 上即使不用前沿裁判模型也比基座高 5.3 分,胜过真值奖励训练与其他 rubric 训练设定。

为什么值得关注

outcome-blind 长程 RL 的信用分配:动态 rubric + 闭式重分布到步骤级 GRPO 优势,AppWorld +15.9 分,无需训练归因模块

English Abstract

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module.

Obsidian 证据

  • 元数据与摘要经 opencli arxiv paper 2609.04094 核对(2026-09-06T04:24:00Z);中文导读锚定摘要陈述的事实与数字。