From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
- ID: a9ce70f4
- 原文链接: https://arxiv.org/abs/2608.27442
- PDF: https://arxiv.org/pdf/2608.27442v1
- 作者: Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
- 日期: 2026-08-27
- 更新: 2026-08-27
- 分类: coding
- 来源类型: paper
- 标签: coding, code-review, benchmark, evaluation, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-29T12:59:16+08:00
中文导读
MCR-Bench(arXiv 2608.27442,cs.SE,ISSTA 2026)把代码评审从单轮静态判定拉回真实的多轮交互场景:这是首个缺陷状态感知的多轮代码评审基准,覆盖 5 种常用编程语言、2,269 个真实多轮评审任务,每个任务带细粒度缺陷元数据(描述/类型/严重度)与跨轮状态标签,刻画缺陷在多轮过程中的完整演化轨迹。主流 LLM 实验给出三点发现:其一,总体能力有限,缺陷检测与缺陷生命周期状态追踪随交互轮数增加显著退化;其二,性能因缺陷类型与严重度大幅波动,语义复杂或低显著度的缺陷更容易漏检;其三,错误机制剖析显示跨轮时间错位与长程记忆不足分别是假阳/假阴的关键驱动。
为什么值得关注
多轮交互下的能力退化曲线对评审类 agent 的上线预期很有参考价值;把假阳归因到跨轮时间错位、假阴归因到长程记忆不足,给出了具体可改进的工程抓手,比笼统的"LLM 评审不可靠"有用得多。
原文(抓取存档·节选)
> Abstract (arXiv 2608.27442, v1 2026-08-27, comment: Accepted at ISSTA 2026)
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi-round interactive nature and the complex problem-solving processes inherent in realistic review scenarios. To bridge this gap, we introduce MCR-Bench, the first defect state-aware benchmark designed for realistic multi-round code review. MCR-Bench covers five commonly-used programming languages and consists of 2,269 real-world multi-round code review tasks, each of which is annotated with fine-grained defect information and cross-round state labels. Each task in MCR-Bench is equipped with fine-grained defect metadata (e.g., description, type, severity) alongside dynamic state annotations, capturing the complete evolutionary trajectory of a defect throughout the multi-round process. We obtain several findings through extensive experiments on MCR-Bench with mainstream LLMs. (1) Limited overall capability: experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases; (2) Defect-sensitive performance: LLMs' performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed; (3) Underlying Failure Mechanisms: our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory.
Obsidian Notes
- 内容由
opencli arxiv paper 2608.27442 -f json拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。