What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
- ID: 00c4ff63
- 原文链接: https://arxiv.org/abs/2608.17719
- PDF: https://arxiv.org/pdf/2608.17719
- 作者: Xiaonan Xu, Wenjing Wu
- 日期: 2026-08-18
- 更新: 2026-08-18
- 分类: cs.SE, cs.AI, cs.CL
- 来源类型: arxiv
- 标签: benchmark, evaluation, llm-api, arxiv, paper
- 质量评分: 4/5
- 抓取时间: 2026-08-20T15:44:49Z
中文导读
依赖商业 LLM API 的系统在厂商废弃旧模型时必须迁移,而迁移决策通常只看聚合基准分。作者在 GPT-5.4 到 GPT-5.6 Sol 的三次两两升级上,对 900 个公开基准题每题每模型查询 50 次,在错误发现率控制与实际显著性阈值下逐题分类为可靠改进、可靠退化、等价或不确定,并用标签置换零模型校准。结果显示九个迁移-基准单元中可靠改进与可靠退化并存:聚合分上涨 7.3 个点的边上仍有 8.3% 的题可靠退化,聚合分下跌的边上也有最多 10.7% 的题可靠改进。指令跟随基准上严格与宽松评分差距在最新迁移拉开 3.9 个点。响应级归档与逐题打分输出已公开。
为什么值得关注
聚合分掩盖题级双向变化:API 迁移验收必须看逐题回归,不能只看总分涨跌
English Abstract
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.
Obsidian 证据
- 来源 digest: 论文流水线 2026-08-20(评分 8.4)。
- 元数据与摘要经 opencli arxiv paper 核对;中文导读锚定摘要陈述的事实与数字。