AI 编程 5.0 · 必读 2026-08-03 · 论文

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

SWE-Touch 把编码 Agent 放到用户会中途修改代码的共享工作区中,用与任务冲突但表面合理的 Counter-Edit 压测 Agent 的状态感知和适应能力摘要报告在 SWE-bench Verified 上平均解决率下降 7.7 个百分点,且问题在更长周期的 SWE-Bench Pro 和 DeepSWE 中仍然存在,指向检测工作区变化协调冲突编辑和针对性测试的下一代基准需求

打开原文回到归档

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

  • ID: 6c376ff4
  • 原文链接: https://arxiv.org/abs/2608.02499
  • PDF: https://arxiv.org/pdf/2608.02499v1
  • 作者: Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu
  • 日期: 2026-08-03
  • 更新: 2026-08-03
  • 分类: coding
  • 来源类型: paper
  • 标签: coding-agents, swe-bench, shared-workspace, evaluation
  • 质量评分: 5/5
  • 抓取时间: 2026-08-05T04:19:03Z

中文导读

SWE-Touch 把编码 Agent 放到用户会中途修改代码的共享工作区中,用与任务冲突但表面合理的 Counter-Edit 压测 Agent 的状态感知和适应能力摘要报告在 SWE-bench Verified 上平均解决率下降 7.7 个百分点,且问题在更长周期的 SWE-Bench Pro 和 DeepSWE 中仍然存在,指向检测工作区变化协调冲突编辑和针对性测试的下一代基准需求

为什么值得关注

SWE-Touch benchmarks coding agents under user code edits in shared workspaces.

Grounded relevance: authors, date, arXiv categories, and abstract claims below; no extra experimental claims beyond the abstract/metadata.

arXiv comment: Preprint. Our code is available at https://github.com/Trae1ounG/SWE-Touch

关键信息

  • 论文标题:SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
  • 作者:Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu
  • arXiv:https://arxiv.org/abs/2608.02499
  • 发布时间:2026-08-03
  • arXiv 分类:cs.SE, cs.AI, cs.CL
  • 关联标签:coding-agents, swe-bench, shared-workspace, evaluation

English Abstract

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to task-relevant code that conflict with task completion. SWE-Touch mines task-critical regions from multiple repair trajectories, uses a separate User Patch Generator to construct the edits, and injects them with contextual user messages when agents reach the relevant code. We evaluate nine coding models on SWE-bench Verified, with additional experiments on longer-horizon tasks from SWE-Bench Pro and DeepSWE. Counter-Edit lowers average resolve rate by 7.7 percentage points on SWE-bench Verified, with degradation also persisting on both longer-horizon benchmarks. Trajectory analysis links these failures to limited awareness of the evolving workspace: agents may retain conflicting code or replace it without sufficiently re-inspecting the repository and validating the revised code with targeted tests. These findings show that strong autonomous performance does not yet ensure the state awareness and adaptive behavior needed for shared-workspace collaboration, and point to detecting workspace changes, reconciling conflicting edits with the task, and verifying the affected behavior as key capabilities for future optimization.

English Summary

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to task-relevant code that conflict with task completion. SWE-Touch mines task-critical regions from multiple repair trajectories, uses a separate User Patch Generator to construct the edits, and injects them with contextual user messages when agents reach the relevant code....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • content-fetcher run; entry id 6c376ff4; arXiv 2608.02499.