Agent 与自动化 4.0 · 优秀 2026-09-04 · 论文

Testing Interchangeability in LLM Agent Teams

2609.05279 (cs.AI/cs.MA,9 月 4 日) 直接测"多智能体系统里同角色可互换"这条默认假设:每个设定 8 队从同一底座独立组 队,跨 10 轮形成私有笔记,然后把角色匹配的成员跨队互换对照组是只打断编排不换人的安慰剂结果:任务分数几乎不变, 但每单位进度的通信开销升 16-63%,Hanabi 上互换后成员表现明显变差

打开原文回到归档

Testing Interchangeability in LLM Agent Teams

Source: https://arxiv.org/abs/2609.05279
Authors: Jianxin Gao, Tianyi Yu, Linna Deng, Runze Li, Zining Wang
Published: 2026-09-04
Categories: cs.AI, cs.MA
PDF: https://arxiv.org/pdf/2609.05279v1

Abstract

Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team's history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.

Key Findings

  • Setup: Eight teams per setting are independently formed from one base model on the same tasks, each agent keeping a private notebook across 10 formation episodes. Then role-matched agents are swapped between teams and measured on held-out tasks.
  • Result vs. placebo: Swaps barely move task score, but raise communication-per-unit-progress by 16–63%. In Hanabi, a swapped agent is even more expensive than an inexperienced one — consistent with interference from conventions learned with a former partner.
  • Who pays for the disruption: In Collab-Overcooked, when the agenda-setter is replaced, most extra communication comes from the agent that *stayed*.
  • Ablations (base model, decoding temperature, formation length): Greedy decoding lowers both swap penalty and inter-team drift; doubling formation history raises both.
  • Takeaway: Agents are more fungible in task outcome than in coordination efficiency; longer formation histories amplify swap costs.

中文概要

本文测试产线多智能体系中的一个隐含假设:“填位某个角色的代理与同能力的代理可以互换”。设计为每个设置独立构建 8 个团队,在 10 个形成轮次中保留私有笔记本;然后交换角色匹配的代理对召回外任务重新评估。结果:对高账价控制组 (仅打乱名单不换人) 相比,交换代理带来的任务得分损失很小,但单位进展所需的沟通量上升 16%–63%;Hanabi 中被交换者甚至比初学者更贵,与其原配对伙伴学到的沟通习惯产生干扰一致。三项剖剖变化(基模、采样温度、形成长度)验证:豌套解码能同时降低交换代价和独立团队散度;加倍形成长度会同时抬升两者。结论:代理在任务结果上更互换,但在协同效率上不互换,且形成史越长交换代价越大。