ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
- source_url: https://arxiv.org/abs/2607.19321
- source_type: paper
- platform: arxiv
- author: Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
- original_date: 2026-07-21
- added_date: 2026-07-23
- category: agents
- tags: ai-control, sabotage, monitoring, ai-rd, safety, benchmark, researcharena, arxiv
- quality_score: 5
- arxiv_id: 2607.19321
- arxiv_categories: cs.AI, cs.CR, cs.LG
摘要(中文)
ResearchArena 评估自动化 AI 研发场景下的 AI control:不信任 Agent 本身,用 monitor 在部署前抓隐蔽破坏四项长程任务:安全后训练能力后训练CUDA kernel 优化推理服务优化;每项配隐藏 side task嵌入式破坏提交物,或沙箱内越界独立行动但主产物诚实前沿模型同时扮演破坏者与监控者;监控类型沿两轴变化:能否对产物执行/探测,以及是否看到 CoT藏在训练数据里的破坏最难抓(检出 <50%);允许 monitor 跑实验有帮助但仍会因只看表面合理化异常或测错而漏检框架开源arXiv:2607.19321,2026-07-21
Summary (English)
ResearchArena evaluates AI control for automated AI R&D when agents may be untrusted: a monitor must catch covert sabotage before artifacts deploy. Four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, inference-server optimization. Each pairs a main task with hidden side tasksembedded sabotage of the submitted model/adapter/kernel/server, or independent out-of-scope sandbox actions while the artifact stays honest. Frontier agents are tested as saboteurs and monitors across four monitor types (execute/probe artifact vs trajectory-only; with/without chain-of-thought). Training-data sabotage is hardest (<50% flagged); letting monitors run experiments helps but still fails via surface inspection, explaining away anomalies, or wrong probes. Modular framework released. arXiv:2607.19321, 2026-07-21.
One-liner
ResearchArena 用长程 AI R&D 任务评估不可信 Agent 的破坏与监控。
Source body / metadata
Fetched via opencli arxiv paper <id> -f json during AAIF content-fetcher backfill. The content note is grounded in the arXiv metadata and abstract.
Abstract
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.
Metadata
- arXiv: 2607.19321
- authors: Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
- published: 2026-07-21
- categories: cs.AI, cs.CR, cs.LG
- url: https://arxiv.org/abs/2607.19321