Why Do Multi-Agent LLM Systems Fail?
- ID: 9eaf642d
- 原文链接: https://arxiv.org/abs/2503.13657
- 作者: Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A....
- 日期: 2025-03-17
- 分类: agents
- 来源类型: paper
- 标签: multi-agent, failure-taxonomy, mast, berkeley, agent-frameworks, evaluation
- 质量评分: 4/5
- 抓取时间: 2026-09-21T23:30Z
中文摘要
arXiv:2503.13657(MAST,NeurIPS 2025)系统标注 7 个主流多智能体框架 1600+ 条执行轨迹,提出首个多智能体失败分类法 MAST3 大类共 14 种失败模式:系统设计问题(FC1,合计约 41%)智能体间失配(FC2,合计约 22%)任务验证失败(FC3,合计约 21%)具体高频模式:步骤重复 FM-1.3 约 15.7%推理与动作不一致 FM-2.6 约 14%不识别任务已完成 FM-1.5 约 12.4%违反规格 FM-1.1 约 11.8%ChatDev 这种带显式验证器的框架比无验证器框架失败更少,但在 ProgramDev 上通过率只有 33.33%证明很多失败不是模型能力不够,而是组织方式不对
English Abstract
arXiv:2503.13657 (MAST, NeurIPS 2025) systematically annotates 1600+ execution traces across 7 popular multi-agent frameworks and proposes the first MAS failure taxonomy MAST. Three categories totaling 14 failure modes: system design issues (FC1, ~41%), inter-agent misalignment (FC2, ~22%), task verification failure (FC3, ~21%). The most-cited concrete modes are step repetition FM-1.3 (~15.7%), reasoning-action mismatch FM-2.6 (~14%), task-completion-blindness FM-1.5 (~12.4%), and spec violation FM-1.1 (~11.8%). Frameworks with explicit verifiers (e.g. ChatDev) fail less than those without, but ChatDev still only achieves 33.33% on ProgramDev direct evidence that many failures stem from mis-organization rather than raw model capability. Inter-annotator agreement kappa=0.88, validating the taxonomy.
为什么值得关注
MAST 标注 7 框架 1600+ 轨迹 14 种失败模式 FC1/FC2/FC3 各占 41/22/21% ChatDev ProgramDev 也只有 33%
Obsidian 证据摘要
来源: OpenClaw定时任务/AK-RSS-Digest(89源精选)/2026-09-21-AK-RSS-Digest(89源精选).md + 对应 evidence-2026-09-21 文件