AI 编程 4.0 · 优秀 2026-09-18 · 论文

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

计算机使用 agent(CUA)沿图形交互与写代码/命令行两条线分别发展,而真实数字工作需要两者交织论文研究混合 CUA:自主决定何时探索界面实现软件运行并以视觉方式验证产物提出 RecreationWorld,围绕'复刻'构建五平台框架(Ubuntu/macOS/Windows/Android/Web):给定一个运行中的参考实现,agent 必须自行发现其行为并构建忠实实现,无预定工作流;统一 harness 同时提供原生 GUI 控制与编码工具运行参考充当隐藏行为测试的 oracle,提供执行接地奖励轨迹生成靠高质量开源应用扩规模;在这些轨迹上训练的模型在五个分布外编码与混合计算机使用基准上均提升,且更频繁地视觉验证自己的渲染产物

打开原文回到归档

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

  • ID: 680726bd
  • 原文链接: https://arxiv.org/abs/2609.22000
  • PDF: https://arxiv.org/pdf/2609.22000v2
  • 作者: Shuai Bai, Jiayong Deng, Sicheng Fan, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou
  • 日期: 2026-09-18
  • 更新: 2026-09-21
  • 分类: coding
  • arXiv 分类: c, s, ., C, L, ,, , c, s, ., S, E
  • 来源类型: paper
  • 标签: computer-use, hybrid-agent, rl-environments, gui, benchmark
  • 质量评分: 4/5
  • 抓取时间: 2026-09-22T04:22:47Z

中文导读

计算机使用 agent(CUA)沿图形交互与写代码/命令行两条线分别发展,而真实数字工作需要两者交织论文研究混合 CUA:自主决定何时探索界面实现软件运行并以视觉方式验证产物提出 RecreationWorld,围绕'复刻'构建五平台框架(Ubuntu/macOS/Windows/Android/Web):给定一个运行中的参考实现,agent 必须自行发现其行为并构建忠实实现,无预定工作流;统一 harness 同时提供原生 GUI 控制与编码工具运行参考充当隐藏行为测试的 oracle,提供执行接地奖励轨迹生成靠高质量开源应用扩规模;在这些轨迹上训练的模型在五个分布外编码与混合计算机使用基准上均提升,且更频繁地视觉验证自己的渲染产物

为什么值得关注

RecreationWorld 五平台混合 CUA 环境:从运行参考逆向复刻行为,隐藏行为测试提供执行接地奖励,OOD 基准全面提分

关键信息

  • 论文标题:RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
  • 作者:Shuai Bai, Jiayong Deng, Sicheng Fan, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou
  • arXiv:https://arxiv.org/abs/2609.22000
  • PDF:https://arxiv.org/pdf/2609.22000v2
  • 发布时间:2026-09-18
  • 最近更新:2026-09-21
  • arXiv 主分类:cs.CL
  • arXiv 全部分类:c, s, ., C, L, ,, , c, s, ., S, E
  • 评论:N/A
  • 关联标签:computer-use, hybrid-agent, rl-environments, gui, benchmark

English Abstract

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards. We scale trajectory generation with high-quality open-source applications. Models trained on these trajectories improve across five out-of-distribution coding and hybrid computer-use benchmarks and more frequently verify their rendered outputs, providing evidence of transfer beyond recreation. For held-out evaluation, we introduce RecreationBench, comprising 250 diverse tasks across domains and platforms. Reference-grounded programmatic and visual assertions cover action-conditioned outcomes at multiple interaction depths; each is validated on the reference and by human reviewers before the suite is frozen for automatic scoring. GPT-6 Astra leads at 58.1% overall, but passes all programmatic tests on just 2.8% of tasks. Agents reproduce static interface structure more reliably than interactions and computed outputs, while generated applications remain smaller and more monolithic than their references. We release the benchmark, environments, and test suites.

English Summary

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。