Agent 与自动化 4.0 · 优秀 2026-09-16 · 论文

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

Flag Game 用玩具模型研究多智能体集体信念形成的机制:隐藏国旗为真值,每个有界 agent 只能看到私有裁剪,通过交换信念权衡同伴社会证据做决策模型复现了丰富的集体现象:性能随群体规模非单调变化社会感知提示与团队多样性带来增益,并发现小规模下的集体信念崩塌会随规模增长转为集体信念极化极化正是大规模性能下降的原因,同时也制造了信念多样性方法上提出 social circuit attribution 预测哪个 agent 的哪种视角对集体动态最关键,并用 agent patching 因果干预验证;对更大群体另建与经验相图吻合的统计力学理论作者称之为 mechanistic swarm interpretability 的第一步

打开原文回到归档
Source: Flag Game: A Toy Model for Mechanistic Swarm Interpretability · platform: arxiv · authors: Elizabeth Pavlova, Hidenori Tanaka · date: 2026-09-16

TL;DR(中文摘要)

AI agent 的涌现协调行为正带来关键安全风险,其核心驱动之一是世界信念的快速形成与传播;要 collective alignment,就需要对这一机制的可解释理解。本文提出 Flag Game:一个研究集体信念形成机制的玩具模型。一面隐藏国旗是 ground truth,每个有界 agent 只能看到私有裁剪区域,但可以交换信念并对同伴的社会证据加权。模型虽简单,却复现了丰富的集体现象:性能随种群规模非单调变化、social-awareness 提示与团队多样性带来准确率增益、组织结构影响强烈。作者识别出:小种群下的集体信念崩塌(belief collapse)随种群增长转变为集体信念极化(polarization);极化导致大种群性能下降,但也制造了集体信念的多样性。机制分析用两条互补路径:一是 social circuit attribution——预测哪个 agent、哪种观点对集体动态影响最大,并用因果干预(agent patching)验证预测,但其效力随种群增大而衰减;二是为更大种群建立统计力学理论,与经验相图吻合。这是迈向 mechanistic swarm interpretability 的第一步。

Summary (English)

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. The Flag Game is a toy model for studying the mechanisms of collective belief formation: a hidden country flag defines the ground truth, and each bounded agent directly observes only a private crop but can exchange beliefs and weigh social evidence from peers. Despite its simplicity, the Flag Game reproduces rich collective phenomenology: non-monotonic scaling of performance with population size, accuracy gains from social-awareness prompting and team diversity, and strong effects of organizational structure. Collective belief collapse at small population sizes turns into collective belief polarization as the population grows; polarization causes the performance decline at large population sizes but creates diversity in collective beliefs. The authors dissect the mechanisms with two complementary approaches: social circuit attribution, which predicts which agent and what view matter most to collective dynamics, verified by causal interventions (agent patching) whose efficacy decreases with population size; and a statistical mechanical theory for larger populations that matches the empirical phase diagram. Together these results take a first step toward mechanistic swarm interpretability.

入库依据

opencli arxiv paper 2609.19124 -f json 抓取 arXiv 元数据(标题/作者/摘要/分类/日期),内容由摘要直接支撑。

补充

arXiv 2609.19124,primary cs.AI,categories: cs.AI, cond-mat.dis-nn, cond-mat.stat-mech, cs.MA, physics.soc-ph;comment: 21 pages, 10 figures;submitted 2026-09-16。