Agent 与自动化 4.0 · 优秀 2026-08-07 · 论文

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

TRIAL 是一种面向代理强化学习的轨迹相对后视蒍istillation 框架,采用统一的轮次对齐评分协议对每个决策轮次,提取其实现后枖的结果视图,并在普通和后视条件下下评估响应,以符号化对数概率差确定 token 级监督方向与强度在 WebShop 和 ALFWorld 上,TRIAL 在八种后端/环境/指标组合中均优于 GRPO;在 WebShop + Qwen3-1.7B 上成功率从 56.4% 提升至 75.2%

打开原文回到归档

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

Source: https://arxiv.org/abs/2608.07371
Author: Haoyu Zheng, Yun Zhu, Qing Wang, Wenqiao Zhang
Date: 2026-08-07
Category: agents
Type: paper
Quality: 4/5
Tags: agentic-rl, hindsight-distillation, grpo, trajectory-relative, arxiv

English Summary

TRIAL is a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol for agentic RL. For each decision turn, it extracts an outcome view of that decision's realized consequence and evaluates the response under ordinary and hindsight-conditioned contexts. The signed log-probability gap determines direction and local strength of token-level supervision, while turn-level magnitudes are normalized jointly over the realized trajectory. Experiments on WebShop and ALFWorld show TRIAL outperforms GRPO across all eight combinations of backbone, environment, and metric. On WebShop with Qwen3-1.7B, TRIAL improves success rate from 56.4% to 75.2% and task score from 78.7% to 85.7%.

中文概要

TRIAL 是一种面向代理强化学习的轨迹相对后视蒍istillation 框架,采用统一的轮次对齐评分协议对每个决策轮次,提取其实现后枖的结果视图,并在普通和后视条件下下评估响应,以符号化对数概率差确定 token 级监督方向与强度在 WebShop 和 ALFWorld 上,TRIAL 在八种后端/环境/指标组合中均优于 GRPO;在 WebShop + Qwen3-1.7B 上成功率从 56.4% 提升至 75.2%

*Added via external-scan on 2026-08-11*