模型与实验室 4.0 · 优秀 2026-09-03 · 论文

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments (arXiv 2609.04...

阿里 Qwen 团队把存量 agent 轨迹变成可复用的可执行环境:轨迹中的工具执行历史暴露了原环境的结构与内容,Terminal-Universe 重放文件操作恢复 agent 修改前的部分工作区,再由补全 agent 补齐缺失文件与依赖;在恢复的工作区上既重建原始意图任务,也合成全新任务,并沿广度(跨工作区跨代码库查询)与深度(经 user agent 扩展为带迭代反馈的多轮会话)两个轴扩展应用于公开终端 agent 轨迹,产出 37.3k 任务充分环境...

打开原文回到归档

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

  • ID: d4af9a81
  • 原文链接: https://arxiv.org/abs/2609.04148
  • PDF: https://arxiv.org/pdf/2609.04148v1
  • 作者: Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang, Dayiheng Liu
  • 日期: 2026-09-03
  • 分类: agents
  • 来源类型: paper
  • 标签: terminal-agents, environments, post-training, trajectory-replay, rl-tasks, qwen, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-05T12:33:07Z

中文导读

阿里 Qwen 团队把存量 agent 轨迹变成可复用的可执行环境:轨迹中的工具执行历史暴露了原环境的结构与内容,Terminal-Universe 重放文件操作恢复 agent 修改前的部分工作区,再由补全 agent 补齐缺失文件与依赖;在恢复的工作区上既重建原始意图任务,也合成全新任务,并沿广度(跨工作区跨代码库)与深度(单轮→多轮会话)两个轴扩展;在 37.3k 环境上对 Qwen3.5-27B SFT,Terminal-Bench 2.1 提升 11.9 点,EvoCode-Bench v2 MT@4 提升 13.8 点。

为什么值得关注

轨迹到环境的转换让历史数据获得二次生命周期,是 terminal agent 后训练数据瓶颈的直接解法。

关键信息

  • 论文标题:Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
  • 作者:Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang, Dayiheng Liu
  • arXiv:https://arxiv.org/abs/2609.04148
  • 发布时间:2026-09-03
  • arXiv 分类:cs.AI, cs.CL(primary: cs.AI)
  • 关联标签:terminal-agents, environments, post-training, trajectory-replay, rl-tasks, qwen, arxiv

English Abstract

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.

English Summary

Terminal-Universe replays file operations recorded in an agent trajectory to restore a partial workspace, then a completion agent supplies missing files and dependencies. On the recovered workspace it reconstructs the original intent task and synthesizes new ones, scaling along breadth (cross-workspace queries) and depth (multi-round sessions via a user agent). Applied to public terminal agent trajectories it yields 37.3k task-sufficient environments; SFT on Qwen3.5-27B improves Terminal-Bench 2.1 single-round by 11.9 points and EvoCode-Bench v2 MT@4 multi-round by 13.8 points.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。