AI 编程 4.0 · 优秀 2026-09-03 · 论文

When Models Edit Too Much: On the Fidelity of Minimal Code Edits (arXiv 2609.04061)

研究代码修复中的 over-editing:模型重写超出修复所需范围基于 400 个 BigCodeBench 问题,向参考解注入受控 AST 级损坏,使每个修复任务都有已知最小补丁前沿模型普遍 over-edit,GPT-5.5 也存在高 Pass@1 与不必要大编辑增加认知复杂度并存的现象一条 preservation 指令显著改善:平均超额 Levenshtein 距离 0.195 降至 0.131,新增认知复杂度降 26.6%,Pass@1 反升 2.3 分;且收益不来自更大推理预算或更大模型后训练阶段,SFT 过拟合于已见损坏模式,RL 给出最佳域外编辑保真与性能保持折中编辑保真被确立为代码修复质量的一个独立可度量可学习维度EMNLP 2026 Main2026-09-03 提交 cs.SE

打开原文回到归档

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

  • ID: d66727b3
  • 原文链接: https://arxiv.org/abs/2609.04061
  • PDF: https://arxiv.org/pdf/2609.04061v1
  • 作者: Tongyao Zhu, Wei Hern Lim, Min-Yen Kan
  • 日期: 2026-09-03
  • 分类: coding
  • 来源类型: paper
  • 标签: over-editing, code-repair, edit-fidelity, rlhf, bigcodebench, emnlp-2026, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-05T12:33:07Z
  • 会议/备注: EMNLP 2026 (Main)

中文导读

研究代码修复中的 over-editing:模型重写超出修复所需范围。基于 400 个 BigCodeBench 问题,向参考解注入受控 AST 级损坏,使每个修复任务都有已知最小补丁。前沿模型普遍 over-edit,GPT-5.5 也存在高 Pass@1 与不必要大编辑、增加认知复杂度并存的现象。一条 preservation 指令同时改善这一行为,平均冘余 Levenshtein 距离从 0.195 降到 0.131,额外认知复杂度下降 26.6%,Pass@1 上升 2.3 点。后训练方面,SFT 对已见损坏模式过拿,RL 在跳出分布的 edit-fidelity 与 Pass@1 保持上最佳。

为什么值得关注

对 coding agent 的 diff 最小化给出可操作结论:一条指令即可同时改善保真与通过率。

关键信息

  • 论文标题:When Models Edit Too Much: On the Fidelity of Minimal Code Edits
  • 作者:Tongyao Zhu, Wei Hern Lim, Min-Yen Kan
  • arXiv:https://arxiv.org/abs/2609.04061
  • 发布时间:2026-09-03
  • arXiv 分类:cs.SE, cs.AI, cs.CL(primary: cs.SE)
  • 关联标签:over-editing, code-repair, edit-fidelity, rlhf, bigcodebench, emnlp-2026, arxiv

English Abstract

Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.

English Summary

The paper studies over-editing in LLM code repair: an evaluation framework built from 400 BigCodeBench problems injects controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Frontier models over-edit widely; a single preservation instruction cuts excess Levenshtein from 0.195 to 0.131, drops added cognitive complexity by 26.6%, and lifts Pass@1 by 2.3 points. SFT overfits to seen corruption patterns; RL gives the best out-of-domain edit-fidelity and Pass@1 trade-off.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。