漏洞定位基准:仓库规模下智能体安全分析能力的测量
Source: arXiv:2609.15939 · Category: cs.CR · Authors: Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi
TL;DR
语言模型智能体越来越多地在完整软件仓库上运作,但现有网络安全评测主要测量“能否检测、重现或修复漏洞”,而不是“能否定位到相关代码”。本文研究漏洞定位:给定一个弱点类型与陆生仓库,要求输出与弱点相关的实现文件。提出 VLoc Bench:涵盖 6 个包生态与 147 个 CWE 类别、源自 290 个仓库的 500 个真实世界漏洞,每个任务以修复前、后两个快照配对。在脆弱快照上,智能体仅获得 CWE 描述与只读终端权限,并必须返回受影响文件;在已修复快照上,则需判断记录的漏洞不再存在。在统一智能体接口下评测 27 个语言模型与 4 个静态分析工具。结果表明:仓库规模下的定位仍难,最强系统仅达 0.229 的 File F1,有 38.4% 的任务无任何被测模型正确定位。另一发现:“能定位漏洞文件”不等同于“修复后表现可靠”,能识别脆弱文件的系统在已修复仓库上仍会报告不叫实的位置。这些结论将漏洞定位确立为仓库规模智能体的一项独立能力,亦提供了研究安全智能体如何检索代码、以及何时应当“不报”的设置。
为什么重要
仓库级别安全智能体的评测缺口是“读代码 + 定位”。VLoc Bench 把这个能力单独拆出来并给出严谨的快照配对设计;最强系统 File F1 仅 0.229、绝大多数任务难以正确定位,对智能体代码审计 / 仓库寻越代码的上限是一个硬数据。
原文摘要(English)
Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness class and an unfamiliar repository, identify the implementation files associated with that weakness. We introduce the Vulnerability Localization Benchmark (VLoc Bench), comprising 500 real world vulnerabilities from 290 repositories across six package ecosystems and 147 CWE categories. Each task pairs repository snapshots immediately before and after a security fix. On the vulnerable snapshot, an agent receives only the CWE description and read-only terminal access and must return the affected files; on the patched snapshot, it must determine that the recorded vulnerability is no longer present. We evaluate 27 language models and four static-analysis tools under a common agent interface. Repository-scale vulnerability localization remains difficult: the strongest system achieves 0.229 File F1, and 38.4% of tasks receive no correct localization from any evaluated model. We further find that stronger localization does not imply reliable behavior after remediation: systems that identify vulnerable files effectively can still report unsupported locations on patched repositories. These results establish vulnerability localization as a distinct repository-scale capability and provide a setting for studying both how security agents search for vulnerable code and when they should refrain from reporting it.
基本信息
| 项 | 值 | |------|------| | 论文 ID | 2609.15939 | | 主分类 | cs.CR | | 发表 | 2026-09-14 | | 更新 | 2026-09-14 | | 作者 | Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi | | 备注 | 29 pages, 6 figures, technical report for VLoc-bench | | PDF | 2609.15939 | | abs | https://arxiv.org/abs/2609.15939 |
参考
- arXiv abs: <https://arxiv.org/abs/2609.15939>
- arXiv PDF: <https://arxiv.org/pdf/2609.15939v1>