Agent-Editing World Model: Rethinking World Modeling for LLM Agents
Source: <https://arxiv.org/abs/2609.28416>
Authors: Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen
Published: 2026-09-23
Categories: cs.CL, cs.AI, cs.LG
arXiv: 2609.28416
Abstract
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from task-state contamination, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the Agent-Editing World Model (AEWM), which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines Action Judge to distinguish Critical, Exploratory, and Noisy decisions with State Revision to edit noisy reasoning-action continuations from the same observed history. EditAct integrates these capabilities with real execution. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2-6.7 points over the strongest baseline. AEWM-RFT, rejection-sampling fine-tuning on verified EditAct trajectories, improves over Self-RFT by 2.2-2.6 points across three domains without online AEWM guidance.
Summary
Proposes AEWM, a world model that edits noisy reasoning-action continuations instead of simulating tool responses. Action Judge classifies decisions as Critical / Exploratory / Noisy (70.5% macro-F1, +10.6 vs strongest frontier baseline); State Revision rewrites history; EditAct integrates both with real execution. Across Search / Terminal / SE, EditAct lifts average scores by 3.2-6.7 points over the strongest baseline; AEWM-RFT gains another 2.2-2.6 points without online AEWM guidance.
摘要
提出 Agent-Editing World Model (AEWM):不再载负模拟工具返回,而是编辑历史中噪声性推理-行动序列。Action Judge 将决策划为 Critical / Exploratory / Noisy,macro-F1 达 70.5%(较最强基线 +10.6);State Revision 修改中间状态;EditAct 将两者与真实执行打通。跨 Search / Terminal / SE 六个基准,平均分数较最强基线 +3.2~+6.7;AEWM-RFT 拒样微调另外 +2.2~+2.6 分,且不依赖在线的 AEWM 指导。