研究与学习 4.0 · 优秀 2026-09-01 · 论文

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

这篇论文反对仅根据 pass/fail 矩阵选择 SWE 代理评测子集的传统做法提出 PTA-IRT,在 Item Response Theory 中加入trajectory 作为特权信息:不仅看正确率,还看过程中探索了哪些文件尝试了哪些编辑走了哪些路径在四个 SWE 基准上,在低校准预算下,PTA-IRT 在分数与排名恢复上均优于现有 IRT 基线代码与数据:

打开原文回到归档

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

  • ID: 78488009
  • 原文链接: https://arxiv.org/abs/2609.01603
  • PDF: https://arxiv.org/pdf/2609.01603v1
  • 作者: Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng
  • 日期: 2026-09-01
  • 更新: 2026-09-01
  • 分类: learning
  • 来源类型: paper
  • 标签: llm-agent, swe-benchmark, evaluation, item-response-theory
  • 质量评分: 4/5
  • 抓取时间: 2026-09-03T04:29:22Z

中文导读

这篇论文反对仅根据 pass/fail 矩阵选择 SWE 代理评测子集的传统做法提出 PTA-IRT,在 Item Response Theory 中加入trajectory 作为特权信息:不仅看正确率,还看过程中探索了哪些文件尝试了哪些编辑走了哪些路径在四个 SWE 基准上,在低校准预算下,PTA-IRT 在分数与排名恢复上均优于现有 IRT 基线代码与数据:

为什么值得关注

这篇论文反对仅根据 pass/fail 矩阵选择 SWE 代理评测子集的传统做法提出 PTA-IRT,在 Item Response Theory 中加入trajectory 作为特权信息:不仅看正确率,还看过程中探索了哪些文件尝试了哪些编辑走了哪些路径在四个 SWE 基准上...

关键信息

  • 论文标题:Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
  • 作者:Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng
  • arXiv:https://arxiv.org/abs/2609.01603
  • 发布时间:2026-09-01
  • arXiv 分类:cs.SE, cs.AI, cs.CL
  • 关联标签:llm-agent, swe-benchmark, evaluation, item-response-theory

English Abstract

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation. Under low calibration budgets, PTA-IRT consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks. Code and data are publicly available at https://github.com/DeepSoftwareAnalytics/PTA-IRT.

English Summary

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。