SWE-Prime: Fewer Trajectories, Better Performance
- ID: 4eb766bf
- 原文链接: https://arxiv.org/abs/2608.27449
- PDF: https://arxiv.org/pdf/2608.27449v1
- 作者: Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng
- 日期: 2026-08-27
- 更新: 2026-08-27
- 分类: coding
- 来源类型: paper
- 标签: coding, fine-tuning, data-selection, swe-bench, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-29T12:59:16+08:00
中文导读
SWE-Prime(arXiv 2608.27449,cs.SE,2026-08-27)处理 SFT 数据筛选中的一个老问题:任务成功不等于高质量监督,成功轨迹里仍可能混有无效、冗余或高风险步骤,直接拿去做 SFT 会引入噪声监督,并诱导模型模仿不良解题行为。论文提出多粒度两阶段筛选:第一阶段在轨迹级按过程质量、结果质量与数据代表性筛出高质量且有代表性的成功轨迹子集;第二阶段把连续步骤聚合成语义段,按对最终解的贡献、可学习性与潜在风险逐段评估。SFT 时所有段保留在序列中以维持上下文,仅被选中的段计入损失计算。在 SWE-Bench Pro 与 SWE-Bench Verified 上,用其选出的 10% 轨迹子集训练即超过全量已解决数据集,相对增益最高分别达 12.2% 与 24.2%。
为什么值得关注
数据筛选比堆量更重要的一次量化验证:10% 精选轨迹反超全量数据,且轨迹级+段级的两阶段流程可以直接搬进自家的 agent SFT 数据管线,段级"保留上下文、只算选中段损失"的做法尤其干净。
原文(抓取存档·节选)
> Abstract (arXiv 2608.27449, v1 2026-08-27, comment: 9 pages, 5 figures)
To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors. Therefore, we propose SWE-Prime, a multi-granularity, two-stage SFT data selection method that progressively filters training data at the trajectory and segment levels. Specifically, the first stage performs trajectory-level screening based on process quality, result quality, and data representativeness, selecting a high-quality and representative subset of successful trajectories. The second stage performs segment-level selection by grouping consecutive steps into semantic segments and assessing each segment based on its contribution to the final solution, learnability, and potential risks. During SFT, all segments remain in the sequence to preserve context, while only selected segments contribute to the loss computation. Experiments on SWE-Bench Pro and SWE-Bench Verified show that training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively.
Obsidian Notes
- 内容由
opencli arxiv paper 2608.27449 -f json拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。