研究与学习 4.0 · 优秀 2026-09-22 · 论文

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

SWE-Serve 评测 agent 能否完成生产级推理工程任务:53 个源自 SGLang 近期真实生产变更的仓库级任务,覆盖六个推理工程方向,每题跑在 CPU 或单张 H100 上,用隐藏的功能/回归测试评分,含端到端 serving 测试与校准性能门限,并配可执行 no-op/oracle 对照对抗性 verifier 审查和闭卷执行来保证评测完整性跨 11 个模型31 种模型-算力配置,最佳配置平均 pass@1 达 75%在 19 个有端到端覆盖的任务上,约三分之一的补丁虽通过本地测试却被 model-serving E2E 测试拒绝本地完成与生产正确性之间存在显著鸿沟

打开原文回到归档

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

  • ID: 779d18aa
  • 原文链接: https://arxiv.org/abs/2609.26777
  • PDF: https://arxiv.org/pdf/2609.26777v1
  • 作者: Jennifer Williams, Dave Farris, Jeff Farris, Jiantao Jiao
  • 发布日期: 2026-09-22
  • 更新日期: 2026-09-22
  • 分类: cs.AI
  • 来源类型: paper
  • 标签: arxiv, benchmark, coding-agent, sglang, inference-serving, paper
  • 质量评分: 4/5
  • 抓取时间: 2026-09-24T12:30:00Z

中文导读

现有基准对生产级推理工程覆盖有限:仓库级 SWE 基准不针对推理,通用 terminal-agent 基准只含少量推理任务,专用推理基准又集中在孤立 kernel 生成或性能优化上SWE-Serve 把这个空白补上:从 SGLang 近期真实生产变更中抽出 53 个仓库级任务,覆盖六个推理工程方向,实现一个推理 feature 往往要协同修改 serving 栈多处——模型支持、runtime 执行、公开 API每题跑在 CPU 或单张 H100 上,用隐藏的功能与回归测试评分,适用时还含端到端 serving 测试和校准过的性能门限;可执行的 no-op/oracle 对照、对抗性 verifier 审查与闭卷执行用来保证评测完整性跨 11 个模型、31 种模型-算力配置,最佳配置平均 pass@1 为 75%关键发现在 19 个有端到端覆盖的任务上:约三分之一的补丁通过了其余所有测试,却被 model-serving E2E 测试拒绝(verifier 口径 45.9%,剔除 E2E 后 69.4%),本地完成任务与生产正确性之间存在可直接度量的鸿沟

为什么值得关注

SGLang 真实生产变更做成 53 题 agent 基准:最强配置 pass@1 75%,但本地通过的补丁约 1/3 被 E2E serving 测试打回。

关键信息

  • 论文标题:SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
  • 作者:Jennifer Williams, Dave Farris, Jeff Farris, Jiantao Jiao
  • arXiv:https://arxiv.org/abs/2609.26777
  • 发布时间:2026-09-22
  • 更新时间:2026-09-22
  • arXiv 分类:cs.AI, cs.SE
  • 关联标签:arxiv, benchmark, coding-agent, sglang, inference-serving, paper

English Abstract

We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited coverage of production inference engineering: repository-level software engineering benchmarks do not target inference, while general terminal-agent benchmarks include only a few inference tasks. Dedicated inference benchmarks, meanwhile, focus primarily on isolated kernel generation or performance optimization rather than repository-scale production feature implementation. SWE-Serve provides 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families. Each task executes on either CPU or a single GPU (H100) and is evaluated with hidden functional and regression tests, including, where applicable, end-to-end (E2E) serving tests and calibrated performance gates. Executable no-op and oracle controls, adversarial verifier review, and closed-book execution support task validity and evaluation integrity. Across 11 models and 31 model-effort configurations, the best-performing configuration achieves 75% mean pass@1. SWE-Serve exposes a substantial gap between completing tasks locally and achieving production correctness. On 19 tasks with end-to-end coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test (45.9% under the verifier versus 69.4% with E2E tests excluded from scoring), with pass rate increasing for each model's best-performing configuration. By making the production correctness gap directly measurable, SWE-Serve enables the field to track whether future agents move beyond completing tasks locally to achieving production correctness.

English Summary

SWE-Serve benchmarks agents on production inference engineering: 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families, each run on CPU or a single H100 and scored with hidden functional and regression tests, E2E serving tests, and calibrated performance gates, with no-op/oracle controls, adversarial verifier review, and closed-book execution for integrity. Across 11 models and 31 model-effort configurations the best reaches 75% mean pass@1. On 19 tasks with E2E coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test, exposing a measurable gap between local task completion and production correctness....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页面属于 content-fetcher 高分条目补全生成,不修改原条目字段,仅产出可读的中英双语页面。