cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
- ID: 2726d6b1
- 原文链接: https://arxiv.org/abs/2609.40284
- PDF: https://arxiv.org/pdf/2609.40284
- 作者: Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh
- 发布时间: 2026-09-30
- 更新: N/A
- 分类: agents
- 来源类型: paper
- 标签: computer-use, benchmark, agent-speed, cost
- 质量评分: 4/5
- 抓取时间: 2026-10-02T04:30:58Z
中文导读
CUA(图形界面操作 agent)在标准基准上已超过人类,但速度与成本评估处于可复现性危机:各基准的机器与容器配置差异混淆执行速度测量cua-speedrun 提供统一虚拟机 + 执行流水线 + 通用 agent 接口,单个 agent 实现可跨 4 个 CUA 基准运行,并测量推理努力harness环境延迟对性能/速度/成本的影响:没有单一模型家族三项全优;开源权重模型不在前沿;反直觉的是对部分模型提高推理努力反而缩短完成时间,更快的环境 IO 也可能拖慢整体完成时间;多数基准可在不损失统计功效的前提下缩减任务集代码与基础设施开源在 cuaspeedrun.com
为什么值得关注
给 CUA 速度/成本测量立标准:统一 VM 与 agent 接口跨 4 基准,揭示推理努力与 IO 延迟的反直觉效应
The abstract describes a uniform virtual-machine setup and common agent interface spanning four CUA benchmarks; findings include no single model family optimal across performance, speed, and cost, no open-weight model on the frontier, more reasoning effort sometimes speeding task completion, and evaluation task sets shrinkable without statistical-power loss.
关键信息
- 论文标题:cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
- 作者:Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh
- arXiv: https://arxiv.org/abs/2609.40284
- 发布时间:2026-09-30
- arXiv 分类:cs.LG, cs.AI, cs.CL
- 关联标签:computer-use, benchmark, agent-speed, cost
英文摘要
Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at this https URL (https://cuaspeedrun.com/).
English Summary
cua-speedrun introduces standardized infrastructure and task sets for evaluating the speed and cost of computer-use agents: a uniform VM setup, a shared execution pipeline, and a common agent interface across four CUA benchmarks. No single model family is optimal on performance, speed, and cost simultaneously; no open-weight model is on the frontier; for some models higher reasoning effort speeds up completion, and faster environment I/O can slow overall completion; most benchmarks can shrink task sets without degrading statistical power. Code, infrastructure, and analysis are open at cuaspeedrun.com.
Obsidian Notes
- 本页为内容补齐(content backfill):条目早已入库,本次补写内容页。
- 内容基于 arXiv abs 页面元数据与摘要生成(opencli arxiv 返回 HTTP 429 后的 abs 页面回退)。
- 中文导读与价值判断均锺定在条目已有摘要与本次拉取的论文摘要上,未添加摘要之外的实验细节。