AI 编程 5.0 · 必读 2026-08-24 · 论文

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

现有 benchmark 只验证行为正确性,不检查迁移是否真的发生,给了 agent 抄旧实现骗过测试的空间(作者称为 Blindness)SWE Refactor Bench 用 20 个整仓迁移任务覆盖 4 类技术债,配合三阶段评测:Migration Audit 确认迁移确实发生固定测试套件验证行为再用 6 个独立 coding agent 生成针对性测试抓隐藏行为差异520 次运行(8 个前沿模型26 组 model-effort 配置)只有 28 次(5.4%)三段全过,20 个任务里 13 个没有任何被接受的解,最好的 claude-opus-5 也只拿 47.0/100作者结论:迁移完整性与行为正确性是两种能力,多数运行要么保行为跳过迁移要么做了迁移弄坏行为,整仓级完美迁移目前做不到

打开原文回到归档

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

  • ID: 4be254ee
  • 原文链接: https://arxiv.org/abs/2608.23564
  • PDF: https://arxiv.org/pdf/2608.23564v1
  • 作者: D, e, y, a, o, , H, o, n, g, ,, , Y, i, z, h, e, , C, h, i, ,, , W, e, n, y, i, , L, i, ,, , X, i, a, o, q, i, u, , W, a, n, g, ,, , M, i, n, g, j, u, , G, a, o, ,, , K, a, i, s, e, n, , Y, a, n, g, ,, , B, i, n, g, x, i, a, n, g, , H, e, ,, , Y, o, u, j, i, e, , Z, h, e, n, g, ,, , C, a, l, v, i, n, , X, i, a, o, ,, , Q, i, n, h, u, a, i, , N, a
  • 日期: 2026-08-24
  • 更新: 2026-08-24
  • 分类: coding
  • 来源类型: paper
  • 标签: swe-refactor-bench, coding-agents, benchmark, technical-debt, long-horizon
  • 质量评分: 5/5
  • 抓取时间: 2026-08-26T04:27:05Z

中文导读

现有 benchmark 只验证行为正确性,不检查迁移是否真的发生,给了 agent 抄旧实现骗过测试的空间(作者称为 Blindness)SWE Refactor Bench 用 20 个整仓迁移任务覆盖 4 类技术债,配合三阶段评测:Migration Audit 确认迁移确实发生固定测试套件验证行为再用 6 个独立 coding agent 生成针对性测试抓隐藏行为差异520 次运行(8 个前沿模型26 组 model-effort 配置)只有 28 次(5.4%)三段全过,20 个任务里 13 个没有任何被接受的解,最好的 claude-opus-5 也只拿 47.0/100作者结论:迁移完整性与行为正确性是两种能力,多数运行要么保行为跳过迁移要么做了迁移弄坏行为,整仓级完美迁移目前做不到

为什么值得关注

整仓迁移三阶段评测:8 个前沿模型 520 次运行仅 5.4% 全通过,最好的 claude-opus-5 也只有 47.0/100

对构建长视野 coding-agent 评测或做选型的人,这套三阶段协议可直接复用:迁移审计先确认迁移真实发生,再跑固定行为测试套件,最后由 6 个独立 coding agent 生成针对性测试查隐藏行为差异,堵住了抄旧实现骗过测试的漏洞。

关键信息

  • 论文标题: SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
  • 作者: D, e, y, a, o, , H, o, n, g, ,, , Y, i, z, h, e, , C, h, i, ,, , W, e, n, y, i, , L, i, ,, , X, i, a, o, q, i, u, , W, a, n, g, ,, , M, i, n, g, j, u, , G, a, o, ,, , K, a, i, s, e, n, , Y, a, n, g, ,, , B, i, n, g, x, i, a, n, g, , H, e, ,, , Y, o, u, j, i, e, , Z, h, e, n, g, ,, , C, a, l, v, i, n, , X, i, a, o, ,, , Q, i, n, h, u, a, i, , N, a
  • arXiv: https://arxiv.org/abs/2608.23564
  • 发布时间: 2026-08-24
  • arXiv 分类: c, s, ., C, L, ,, , c, s, ., A, I, ,, , c, s, ., S, E
  • 备注: N/A
  • 关联标签: swe-refactor-bench, coding-agents, benchmark, technical-debt, long-horizon

English Abstract

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs ($5.4\%$) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores $47.0/100$. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, $58\%$ reach $99\%$ of the fixed checks, yet only $26\%$ reach $100\%$. Agent capability differs across migration categories: agents score $31.4$ on build toolchain rewrites but only $5.6$ on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

English Summary

Existing benchmarks evaluate only behavioural correctness, not whether a migration actually occurred, letting agents copy the original implementation to pass tests (the authors call this Blindness). SWE Refactor Bench contributes 20 whole-repository migrations covering 4 kinds of technical debt, scored by a three-stage protocol: Migration Audit verifies the migration occurred, Behavioural Tests measure correctness with a fixed suite, and Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 runs (5.4%) pass all three stages, 13 of 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness appear to be distinct abilities.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页由 AAIF content-fetcher 定时任务生成(2026-08-26),仅新增内容页,未改动 entries.json。