Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
- ID: 78488009
- 原文链接: https://arxiv.org/abs/2609.01603
- PDF: https://arxiv.org/pdf/2609.01603v1
- 作者: Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng
- 日期: 2026-09-01
- 更新: 2026-09-01
- 分类: learning
- 来源类型: paper
- 标签: llm-agent, swe-benchmark, evaluation, item-response-theory
- 质量评分: 4/5
- 抓取时间: 2026-09-03T04:29:22Z
中文导读
这篇论文反对仅根据 pass/fail 矩阵选择 SWE 代理评测子集的传统做法提出 PTA-IRT,在 Item Response Theory 中加入trajectory 作为特权信息:不仅看正确率,还看过程中探索了哪些文件尝试了哪些编辑走了哪些路径在四个 SWE 基准上,在低校准预算下,PTA-IRT 在分数与排名恢复上均优于现有 IRT 基线代码与数据:
为什么值得关注
这篇论文反对仅根据 pass/fail 矩阵选择 SWE 代理评测子集的传统做法提出 PTA-IRT,在 Item Response Theory 中加入trajectory 作为特权信息:不仅看正确率,还看过程中探索了哪些文件尝试了哪些编辑走了哪些路径在四个 SWE 基准上...
关键信息
- 论文标题:Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
- 作者:Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng
- arXiv:https://arxiv.org/abs/2609.01603
- 发布时间:2026-09-01
- arXiv 分类:cs.SE, cs.AI, cs.CL
- 关联标签:llm-agent, swe-benchmark, evaluation, item-response-theory
English Abstract
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation. Under low calibration budgets, PTA-IRT consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks. Code and data are publicly available at https://github.com/DeepSoftwareAnalytics/PTA-IRT.
English Summary
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。