AI 编程 4.0 · 优秀 2026-09-03 · 论文

PatchBench: Evaluating AI Agents for Vulnerability Patching

自动漏洞修复的评测有效性研究:作者提出补丁相似度度量,检测智能体是否复现了记忆中的历史开发者补丁平均 25% 的智能体补丁与历史补丁高度相似,说明记忆化是此类评测的真实有效性威胁;同时智能体常利用基准结构,在崩溃栈上打补丁压制崩溃而非修根因,以通过 PoC 验证结论:只测 PoC 是否不再崩溃的评测方式不足以证明修复能力

打开原文回到归档

PatchBench: Evaluating AI Agents for Vulnerability Patching

  • ID: fa38e2d2
  • 原文链接: https://arxiv.org/abs/2609.04075
  • PDF: https://arxiv.org/pdf/2609.04075v1
  • 作者: Chihao Shen, Jiacheng Li, Aastha Mahajan, Jeffery Siyuan Tian, Yonghwi Kwon, Yizheng Chen
  • 日期: 2026-09-03
  • 更新: 2026-09-03
  • 分类: cs.CR, cs.AI, cs.SE
  • 来源类型: arxiv
  • 标签: vulnerability-patching, benchmark, memorization, evaluation-validity, security, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-06T04:24:00Z

中文导读

现有漏洞修复评测通常只验证 PoC 输入是否还触发崩溃,作者指出这留下两类有效性威胁:智能体可能复现记忆中的历史开发者补丁,也可能生成只压制崩溃的表面修复。在 C/C++ 漏洞修复任务上,作者提出补丁相似度度量,发现平均 25% 的智能体补丁与历史开发者补丁高度相似,补丁记忆化是此类评测的真实威胁;同时智能体常利用基准结构,直接在崩溃栈上打补丁压制崩溃,而非定位并修复漏洞根因。PatchBench 只选取真值修复位于崩溃栈之外的漏洞,并用漏洞移植与代码变异把历史漏洞迁移进新仓库上下文,降低表面修复与记忆化风险,同时引入同时评估安全性与语义正确性的补丁验证方法。在包括 AIxCC 前三名在内的 11 个 SOTA 智能体上,仅用 PoC 验证平均会把修复解决率虚高 1.83 倍。

为什么值得关注

漏洞修复评测的两个有效性漏洞:25% 补丁是背出来的历史补丁,还有栈上压崩溃骗过 PoC只测崩溃消失不算会修漏洞

English Abstract

AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash. This leaves two key threats to validity: agents may reproduce memorized historical developer patches, or they may generate surface-level fixes that only suppress the reported crash. We study these concerns for C/C++ vulnerability patching. We introduce a patch similarity metric to detect memorized patches. On average, 25% of the agent patches exhibit substantial similarity to historical developer patches, indicating that patch memorization is a real threat to the validity of vulnerability patching evaluations.

Obsidian 证据

  • 元数据与摘要经 opencli arxiv paper 2609.04075 核对(2026-09-06T04:24:00Z);中文导读锚定摘要陈述的事实与数字。