One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
- ID: 926fb43d
- 原文链接: https://arxiv.org/abs/2608.12253
- PDF: https://arxiv.org/pdf/2608.12253v1
- 作者: Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
- 日期: 2026-08-12
- 更新: 2026-08-12
- 分类: agents
- 来源类型: paper
- 标签: multi-agent, rl, simulator-collapse, human-ai-interaction, cs.cl, cs.cl-cs.ai-cs.lg
- 质量评分: 4/5
- 抓取时间: 2026-08-14T04:19:48Z
中文导读
论文指出多代理 RL 中常被忽视的模拟器坍塑:用一个模式跳起的 LLM 模拟用户,会让策略过拟合该模式,转移到未见模拟器或真实用户时性能崩坏提出理论框架并评估了多代理 RL 设计在人机交互上的边界
为什么值得关注
多代理 RL 中的单 LLM 用户模拟器会坍塑,让策略过拟合现有模式
关键信息
- 论文标题:One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
- 作者:Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
- arXiv:https://arxiv.org/abs/2608.12253
- 发布时间:2026-08-12
- arXiv 分类:cs.CL, cs.AI, cs.LG
- 关联标签:multi-agent, rl, simulator-collapse, human-ai-interaction, cs.cl, cs.cl-cs.ai-cs.lg
- 论文备注:41 pages, 28 figures
English Abstract
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.
English Summary
The paper shows that single-LLM user simulators in multi-agent RL are mode-collapsed, so policies trained against them overfit to a narrow strategy and transfer poorly to unseen simulators and real users. It formalizes simulator collapse and reports the implications for human-AI interaction benchmarks that rely on a single frozen simulator LLM.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。