Discriminative World Models for Web Agents
- ID: da3f4917
- 原文链接: https://arxiv.org/abs/2609.02885
- PDF: https://arxiv.org/pdf/2609.02885v1
- 作者: Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig
- 日期: 2026-09-02
- 更新: 2026-09-02
- 分类: agents
- 来源类型: paper
- 标签: web-agents, world-models, process-reward-model, test-time-compute
- 质量评分: 4/5
- 抓取时间: 2026-09-04T04:22:30Z
中文导读
论文针对网页 Agent 测试时动作选择中的世界模型训练错配问题:现有世界模型用监督式下一状态预测(生成 HTML/AXTree 快照)训练,但下游 ranker/PRM 依赖的是预测状态在候选动作间的可区分性作者提出 predicted-state matching 训练目标预测表示必须把真实后续状态与备选动作到达的状态区分开并用 WebArena Go-Browse 轨迹构建的分支数据集训练在自建 benchmark 上超过监督式下一状态预测基线,在 WebPRMBench 上提升 PRM 动作排序,在 WebArena-Lite 上提升端到端任务成功率(cs.AI, cs.LG,2026-09-02)
为什么值得关注
论文针对网页 Agent 测试时动作选择中的世界模型训练错配问题:现有世界模型用监督式下一状态预测(生成 HTML/AXTree 快照)训练. 该工作由 arXiv 摘要直接背书:结论、实验设置与指标均出自摘要原文(2026-09-02 提交,cs.AI, cs.LG),适合关注 agents 方向的近期进展。
关键信息
- 论文标题:Discriminative World Models for Web Agents
- 作者:Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig
- arXiv:https://arxiv.org/abs/2609.02885
- 发布时间:2026-09-02
- arXiv 分类:cs.AI, cs.LG
- 关联标签:web-agents, world-models, process-reward-model, test-time-compute
English Abstract
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates to accurately score them. To address this, we introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions. We train these models using a branching web-agent dataset derived from WebArena Go-Browse trajectories, where every decision point contains multiple alternative actions and their resulting states. Experiments on our held-out predicted-state matching benchmark show that our approach outperforms world models trained with supervised next-state prediction. We further show that our approach improves PRM-style action ranking on WebPRMBench compared with action-only PRMs and PRMs augmented with supervised-next-state world models. Finally, on WebArena-Lite, using our world model for test-time action selection improves end-to-end task success. Our project page is available at: https://dhruvpendharkar.github.io/dwm/.
English Summary
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates. The authors introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions, trained on a branching web-agent dataset derived from WebArena Go-Browse trajectories....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。