BEARCUBS: A benchmark for computer-using web agents
Source: https://arxiv.org/abs/2503.07919 Authors: Yixiao Song, Katherine Thai, Chau Minh Pham, Yapei Chang, Mazin Nadaf, Mohit Iyyer Published: 2025-03-10 Updated: 2025-07-24 Categories: cs.AI, cs.CL, cs.LG
中文摘要
BEARCUBS 是一个针对 computer-using web agents 的基准, 包含 111 个信息查找问题. 它要求代理访问实时 Web 内容, 并完成视频理解, 3D 导航等多模态交互. 摘要中的实验结果显示 ChatGPT Agent 达到 65.8% 准确率, 人类准确率为 84.7%, 差距主要来自精细控制, 复杂过滤和执行速度.
英文摘要(Abstract)
Modern web agents possess computer use abilities that allow them to interact with webpages by sending commands to a virtual keyboard and mouse. While such agents have considerable potential to assist human users with complex tasks, evaluating their capabilities in real-world settings poses a major challenge. To this end, we introduce BEARCUBS, a "smallbut mighty" benchmark of 111 information-seeking questions designed to evaluate a web agent's ability to search, browse, and identify factual information from the web. Unlike prior web agent benchmarks, solving BEARCUBS requires (1) accessing live web content rather than synthetic or simulated pages, which captures the unpredictability of real-world web interactions; and (2) performing a broad range of multimodal interactions (e.g., video understanding, 3D navigation) that cannot be bypassed via text-based workarounds. Each question in BEARCUBS has a corresponding short, unambiguous answer and a human-validated browsing trajectory, allowing for transparent evaluation of agent performance and strategies. A human study confirms that BEARCUBS questions are solvable but non-trivial (84.7% human accuracy), revealing domain knowledge gaps and overlooked details as common failure points. We find that ChatGPT Agent significantly outperforms other computer-using agents with an overall accuracy of 65.8% (compared to e.g., Operator's 23.4%), showcasing substantial progress in tasks involving real computer use, such as playing web games and navigating 3D environments. Nevertheless, closing the gap to human performance requires improvements in areas like fine control, complex data filtering, and execution speed. To facilitate future research, BEARCUBS will be updated periodically to replace invalid or contaminated questions, keeping the benchmark fresh for future generations of web agents.
一句评点
BEARCUBS 用实时 Web 和多模态交互测量电脑使用代理的真实能力差距.
来源与元数据
- arXiv: https://arxiv.org/abs/2503.07919
- PDF: https://arxiv.org/pdf/2503.07919v3
- Authors: Yixiao Song, Katherine Thai, Chau Minh Pham, Yapei Chang, Mazin Nadaf, Mohit Iyyer
- Published: 2025-03-10
- Updated: 2025-07-24
- Categories: cs.AI, cs.CL, cs.LG
- Comments: 16 pages
本文件由 content-fetcher 通过 opencli arxiv paper 获取论文元数据和摘要后回填,不包含未读取全文的额外推断。