Source: Flag Game: A Toy Model for Mechanistic Swarm Interpretability · platform: arxiv · authors: Elizabeth Pavlova, Hidenori Tanaka · date: 2026-09-16
TL;DR(中文摘要)
AI agent 的涌现协调行为正带来关键安全风险,其核心驱动之一是世界信念的快速形成与传播;要 collective alignment,就需要对这一机制的可解释理解。本文提出 Flag Game:一个研究集体信念形成机制的玩具模型。一面隐藏国旗是 ground truth,每个有界 agent 只能看到私有裁剪区域,但可以交换信念并对同伴的社会证据加权。模型虽简单,却复现了丰富的集体现象:性能随种群规模非单调变化、social-awareness 提示与团队多样性带来准确率增益、组织结构影响强烈。作者识别出:小种群下的集体信念崩塌(belief collapse)随种群增长转变为集体信念极化(polarization);极化导致大种群性能下降,但也制造了集体信念的多样性。机制分析用两条互补路径:一是 social circuit attribution——预测哪个 agent、哪种观点对集体动态影响最大,并用因果干预(agent patching)验证预测,但其效力随种群增大而衰减;二是为更大种群建立统计力学理论,与经验相图吻合。这是迈向 mechanistic swarm interpretability 的第一步。
Summary (English)
Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. The Flag Game is a toy model for studying the mechanisms of collective belief formation: a hidden country flag defines the ground truth, and each bounded agent directly observes only a private crop but can exchange beliefs and weigh social evidence from peers. Despite its simplicity, the Flag Game reproduces rich collective phenomenology: non-monotonic scaling of performance with population size, accuracy gains from social-awareness prompting and team diversity, and strong effects of organizational structure. Collective belief collapse at small population sizes turns into collective belief polarization as the population grows; polarization causes the performance decline at large population sizes but creates diversity in collective beliefs. The authors dissect the mechanisms with two complementary approaches: social circuit attribution, which predicts which agent and what view matter most to collective dynamics, verified by causal interventions (agent patching) whose efficacy decreases with population size; and a statistical mechanical theory for larger populations that matches the empirical phase diagram. Together these results take a first step toward mechanistic swarm interpretability.
入库依据
opencli arxiv paper 2609.19124 -f json 抓取 arXiv 元数据(标题/作者/摘要/分类/日期),内容由摘要直接支撑。
补充
arXiv 2609.19124,primary cs.AI,categories: cs.AI, cond-mat.dis-nn, cond-mat.stat-mech, cs.MA, physics.soc-ph;comment: 21 pages, 10 figures;submitted 2026-09-16。