Agent 与自动化 4.0 · 优秀 2026-08-06 · 论文

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

Agent 基准通常只测任务完成度把资源使用当附属统计,而实际部署中本地查询广域搜索复合研究工具更强模型人工升级之间的选择本身就是任务的一部分EcoAgent-Bench 的 304 个真实改编任务(来自 GAIA/HotpotQA/MuSiQue 的五个族)都带定价动作与显式预算,考察四类决策:避免不必要升级证据不足时升级选择模型档位在不成立前提上停止评测了 tool-API 与 workspace-CLI 两种形态的 7 个 LLM agent 及 4 个脚本 oracle...

打开原文回到归档

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

  • ID: 75d7a431
  • 原文链接: https://arxiv.org/abs/2608.05519
  • PDF: https://arxiv.org/pdf/2608.05519v1
  • 作者: Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
  • 日期: 2026-08-06
  • 更新: 2026-08-06
  • 分类: agents
  • 来源类型: paper
  • 标签: benchmark, economic-decision, budget-constraint, agent-evaluation, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-18T05:28:55Z

中文导读

Agent 基准通常只测任务完成度把资源使用当附属统计,而实际部署中本地查询广域搜索复合研究工具更强模型人工升级之间的选择本身就是任务的一部分EcoAgent-Bench 的 304 个真实改编任务(来自 GAIA/HotpotQA/MuSiQue 的五个族)都带定价动作与显式预算,考察四类决策:避免不必要升级证据不足时升级选择模型档位在不成立前提上停止评测了 tool-API 与 workspace-CLI 两种形态的 7 个 LLM agent 及 4 个脚本 oracle;微观平均准确率会奖励一边倒策略(永远升级的控制组微观成功率高但通不过省钱型任务),因此另报经济一致性分数Tool-API agent 微观严格成功率仅 3.9-24.0%(经济一致性至多 7.3%),预算阈值扫描只把 GPT-5.4 的升级率从 0% 拉到 3%:预算内完成与经济动作选择是两种不同能力

为什么值得关注

带定价动作与显式预算的 304 任务基准:预算内完成与经济动作选择是两种不同的能力

该论文发表于 2026-08-06,作者为 Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao,arXiv 分类 cs.AI, cs.CL, cs.LG;以上判断基于论文摘要所述内容。

关键信息

  • 论文标题: EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
  • 作者: Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
  • arXiv: https://arxiv.org/abs/2608.05519
  • 发布时间: 2026-08-06
  • arXiv 分类: cs.AI, cs.CL, cs.LG
  • 备注: 8 pages, 3 figures, 4 tables. Benchmark, dataset (304 budget-conditioned agent tasks), and evaluation harness; artifacts to be released
  • 关联标签: benchmark, economic-decision, budget-constraint, agent-evaluation, arxiv

English Abstract

Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4's escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.

English Summary

Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic, yet in deployment the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. EcoAgent-Bench specifies priced actions and an explicit budget for each of 304 real-derived tasks (five families adapted from GAIA, HotpotQA, and MuSiQue), testing four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. Seven LLM agents in tool-API and workspace-CLI settings plus four oracle scripted controls are evaluated. Micro-averaged accuracy rewards one-sided policies (always-escalate controls score high while failing save-oriented tasks), so the authors also report an economic-consistency score....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。