Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
- ID: ac687980
- 原文链接: https://arxiv.org/abs/2608.12895
- PDF: https://arxiv.org/pdf/2608.12895v1
- 作者: Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
- 日期: 2026-08-13
- 更新: N/A
- 分类: agents
- 来源类型: paper
- 标签: multiagent, reliability, redundancy, correlated-failure, evaluation
- 质量评分: 4/5
- 抓取时间: 2026-08-17T23:51:34+08:00
中文导读
多 agent 系统的组合可靠性界把组件可靠度相乘,其前提——条件独立——被普遍陈述却极少检验。这篇直接检验:同一模型的两个实例做两 agent 交接,在任一失败的任务上 90.0% 共同失败(log OR 6.66,95% CI [6.38, 7.00];phi 0.916);预注册评估 18,000 个任务,全部由确定性代码评分、无 LLM 裁判。换不同模型后六个对比全部降低关联。结论:冗余设计要换模型,别复制自己。
为什么值得关注
同模型双 agent 冗余被系统性高估:90% 共同失败、phi 0.916,冗余必须换模型。
收录理由:用预注册+确定性评分的严格设计检验多 agent 可靠性设计的隐含假设,结论可直接指导冗余架构
Abstract
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.
元数据
- arXiv ID: 2608.12895
- 主分类: cs.AI
- 分类: cs.AI, cs.MA
- 评论: 49 pages, 12 tables, 25 numbered definitions, 18 theorems with full proofs, six experiments, 65 references. Code, analysis scripts, and preregistration: https://github.com/qualixar/agentassert-abc
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.12895(2026-08-17)。