Agent 与自动化 4.0 · 优秀 2026-09-25 · 论文

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

arXiv 2609.31563 把 Steiner 群体任务分类学引入 multi-agent LLM 扩展性分析:在 13 个开源权重模型最多 30 个 agent 的团队上,析取任务(disjunctive)里至少一个 agent 对的概率随团队规模涨 5-20 个点,但直接答案上的多数投票几乎拿不到这个潜力;多轮修订能显著提准确率,可 1 个 peer 和 29 个 peer 的增益几乎一样Fermi 估计这类补偿任务上扩展几乎无益同一模型样本共享的 item 级偏差占平方误差约 87%,平均只降约 6% 误差任务结构 + 输出组合机制才是团队扩展的决定因素

打开原文回到归档

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

原文链接: https://arxiv.org/abs/2609.31563
作者: Carolina Fortuna, Blaz Bertalanic
发布时间: 2026-09-25
源: arXiv外部扫描 (2026-09-29)

摘要

arXiv 2609.31563 把 Steiner 群体任务分类学引入 multi-agent LLM 扩展性分析:在 13 个开源权重模型最多 30 个 agent 的团队上,析取任务(disjunctive)里至少一个 agent 对的概率随团队规模涨 5-20 个点,但直接答案上的多数投票几乎拿不到这个潜力;多轮修订能显著提准确率,可 1 个 peer 和 29 个 peer 的增益几乎一样Fermi 估计这类补偿任务上扩展几乎无益同一模型样本共享的 item 级偏差占平方误差约 87%,平均只降约 6% 误差任务结构 + 输出组合机制才是团队扩展的决定因素

English Summary

Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior....

为什么值得关注

扩展 AAIF 对应主题线

信息源