Agent 与自动化 5.0 · 必读 2026-09-17 · 论文

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

ActObs:在 SFT 阶段不只监督 action token,也监督 trajectory 中已存在的 observation token模型推理时并不生成 observation,但联合监督迫使策略学会建模"动作后果"而不增加参数 / 数据 / 步数SFT 后两者表现相近,GRPO 后差距打开:Qwen3-4B 在 Terminal-Bench 2.0 各 pass@k 预算上都更高;Qwen3-8B 在 pass@1 上略退但 pass@16 +3.4 pp解出更多不同任务;跨域的 aider-polyglot(4B)pass@1 +4.2 pp论文 29 页 / 9 图 / 11 表,2026-09-17 提交 cs.LG

打开原文回到归档

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

  • ID: 3975651b
  • 原文链接: https://arxiv.org/abs/2609.20715
  • PDF: https://arxiv.org/pdf/2609.20715
  • 作者: Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah
  • 日期: 2026-09-17
  • 更新: 2026-09-17
  • 分类: agents
  • 来源类型: paper
  • 标签: arxiv, paper, agentic-rl, sft, grpo, exploration
  • 质量评分: 5/5
  • 抓取时间: 2026-09-21T04:30Z

中文导读

  • ActObs 质疑 SFT 的一个默认习惯:loss 只压在 agent 自己生成的 action token 上,trajectory 里已经存在的 observation token 只当上下文、不当预测目标。
  • 改法本身很轻:把 observation token 也纳入监督(即 ActObs),不加数据、不加参数、不加序列长度、不加 forward pass——部署时 agent 依旧不生成 observation。
  • 差异在 RL 阶段才打开:SFT 后两者接近,GRPO 之后 Qwen3-4B 在 Terminal-Bench 2.0 上每个 pass@k 采样预算都更高;Qwen3-8B 用一点 pass@1 换 pass@16 +3.4 pp 且解出更多不同任务;跨域的 aider-polyglot(4B)pass@1 +4.2 pp。
  • 作者把根因追溯到 SFT:action 与 observation 梯度很快正交,action-only 训练留下大的残余 observation 梯度,环境预测能力反而掉到 base 模型之下;联合监督保住“动作后果”建模,为下游探索留了更好的初始化。

为什么值得关注

ActObs 把 observation token 一起进 SFT loss:零额外成本,让 GRPO 在 Terminal-Bench 与 aider-polyglot 上拿到更高 pass@k

关键信息

  • 论文标题:Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
  • 作者:Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah
  • arXiv:https://arxiv.org/abs/2609.20715
  • 发布时间:2026-09-17
  • arXiv 分类:cs.LG, cs.AI, cs.CL
  • 关联标签:arxiv, paper, agentic-rl, sft, grpo, exploration

English Abstract

Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

English Summary

Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

Obsidian Notes

  • 内容由 opencli arxiv paper 2609.20715 -f json 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。