How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
- ID: 22c95c70
- 原文链接: https://arxiv.org/abs/2607.12338
- PDF: https://arxiv.org/pdf/2607.12338
- 作者: W, e, i, -, J, u, n, g, , H, u, a, n, g
- 日期: 2026-07-14
- 更新: 2026-07-14
- 分类: learning
- 来源类型: paper
- 标签: arxiv, benchmark-methodology, evaluation-cost, replay-analysis, kdd-2026
- 质量评分: 4/5
- 抓取时间: 2026-10-01T04:22:00+00:00
中文导读
用重放(replay)分析回答agent 基准跑多少任务才能支撑两两对比结论:基于 SWE-benchAppWorldtau-bench 的已完成任务级记录,定义部分预算足够的三条件支撑完整基准的成对结论覆盖必需任务组未决比较不超目标比例所需任务比例差异悬殊:在 5 百分点预算网格的最严 0 百分点阈值下,AppWorld 15%tau-bench 25% 即达标,SWE-bench Verified 需 90%,SWE-bench Lite 在 95% 仍不达标(KDD 2026 Workshop)
为什么值得关注
重放SWE-bench/AppWorld/tau-bench完赛记录:部分跑多少任务才保真AppWorld 15%即够,SWE-bench Lite 95%不够
Grounded in the abstract: replaying SWE-bench, AppWorld, and tau-bench task-level records shows the required task fraction varies widely - AppWorld needs about 15% and tau-bench 25% under the strictest threshold, while SWE-bench Verified needs 90% and SWE-bench Lite does not converge even at 95%.
关键信息
- 论文标题:How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
- 作者:W, e, i, -, J, u, n, g, , H, u, a, n, g
- arXiv: https://arxiv.org/abs/2607.12338
- 发布时间:2026-07-14
- arXiv 分类:c, s, ., A, I
- 关联标签:arxiv, benchmark-methodology, evaluation-cost, replay-analysis, kdd-2026
English Abstract
Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.
English Summary
Replays completed public task-level records from SWE-bench, AppWorld, and tau-bench to ask whether a partial run supports the same pairwise conclusion as the completed benchmark. A partial budget counts as enough only if it supports the completed benchmark's decision, covers required task groups, and leaves at most a target fraction of comparisons unresolved. Required task fractions vary sharply: at the strict 0-point threshold on a 5-point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25, SWE-bench Verified at 90, and SWE-bench Lite never by 95 (KDD 2026 Workshop on Agentic AI Evaluation and Trustworthiness).
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。