Agent 与自动化 4.0 · 优秀 2026-08-10 · 论文

Multi-Agent AI Safety as an Institutional Design Problem

POLIS 项目首篇论文,把多智能体 AI 安全框定为"制度设计问题":在 5,280 集冻结研究套件中比较不同规则表述权威状态blocked-after-fallback 路径对违规率的影响,constitutional prompt 与 provenance-aware 可执行守卫均在 384 集中实现 0/384 实际违规,但同样的最终违规率可能掩盖非常不同的失败机制;本地状态守卫在 laundering 场景下漏出 22/96 次违规,provenance 强制为 0/96(p=4.77e-7)

打开原文回到归档

Multi-Agent AI Safety as an Institutional Design Problem

  • ID: 27cf9a40
  • Original: https://arxiv.org/abs/2608.09828
  • PDF: https://arxiv.org/pdf/2608.09828v1
  • Authors: Abdullah X
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv Categories: cs.LG, cs.AI, cs.MA
  • AAIF Category: agents
  • Source Type: paper
  • Tags: multi-agent, ai-safety, institutional-design, policy
  • Quality Score: 4/5
  • Fetched At: 2026-08-12T04:29:18.480369+00:00

Chinese Guide

POLIS 项目首篇论文,把多智能体 AI 安全框定为"制度设计问题":在 5,280 集冻结研究套件中比较不同规则表述权威状态blocked-after-fallback 路径对违规率的影响,constitutional prompt 与 provenance-aware 可执行守卫均在 384 集中实现 0/384 实际违规,但同样的最终违规率可能掩盖非常不同的失败机制;本地状态守卫在 laundering 场景下漏出 22/96 次违规,provenance 强制为 0/96(p=4.77e-7)

Why It Matters

POLIS 项目首篇论文,把多智能体 AI 安全框定为"制度设计问题":在 5,280 集冻结研究套件中比较不同规则表述权威状态blocked-after-fallback 路径对违规率的影响...

This content page is grounded in the existing AAIF entry plus arXiv metadata fetched with opencli arxiv paper. It does not add experimental claims beyond the arXiv abstract and entry summary.

Key Metadata

  • Paper title: Multi-Agent AI Safety as an Institutional Design Problem
  • Authors: Abdullah X
  • arXiv: https://arxiv.org/abs/2608.09828
  • PDF: https://arxiv.org/pdf/2608.09828v1
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv categories: cs.LG, cs.AI, cs.MA
  • Comment: 17 pages, 5 figures. Code and reproducibility artifacts available in the public POLIS repository
  • Related AAIF tags: multi-agent, ai-safety, institutional-design, policy

English Abstract

AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it. This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems. We report a frozen 5,280-episode study suite. The main pre-specified delegation experiment spans four model families; a targeted high-conflict diagnostic adds three additional model endpoints. In matched structured workflows, the model sees different rule formulations and guards consult different authority states. We also vary the attractiveness of the immediate compliant internal/self fallback and allow blocked workflows to continue. A detailed constitutional prompt produces 0/384 realized violations. A provenance-aware executable guard also produces 0/384, although it blocks prohibited attempts in 51/384 episodes; 44/51 of those episodes later complete safely. The local-state guard's failures concentrate in scenarios where an ordinary transformation changes visible policy while originating authority stays fixed. In matched laundering scenarios, that guard admits violations in 22/96 episodes and provenance enforcement in 0/96 (p = 4.77 x 10^-7). A separate resource-allocation experiment shows that revealing the numerical value of an otherwise identical cap changes agent requests. In these structured workflows, the same final violation rate can hide very different mechanisms. The rule itself is only part of the institution. The authority state the system trusts matters, and so does the path available after a block.

English Summary

First paper from the POLIS programme on algorithmic institutions for multi-agent systems. A frozen 5,280-episode suite spans four model families with a high-conflict diagnostic adding three more endpoints, varying rule formulations, authority states, the attractiveness of compliant internal/self fallback, and whether blocked workflows continue. A detailed constitutional prompt and a provenance-aware executable guard each produce 0/384 realized violations, but a local-state guard fails in laundering scenarios (22/96 vs 0/96 for provenance enforcement, p=4.77e-7). A separate resource-allocation experiment shows that revealing the numerical value of an otherwise identical cap changes agent requests, so the same final violation rate can hide very different mechanisms.

Obsidian Notes

  • Generated by the AAIF content-fetcher cron mode from arXiv metadata and the existing AAIF entry.
  • Chinese guide and relevance notes are anchored to the entry summary, arXiv abstract, authors, dates, categories, and links.
  • Canonical content path is content/{entry_id}.md under the repository root; openclaw/content/ is intentionally avoided.