Agent 与自动化 4.0 · 优秀 2026-09-18 · 论文

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

生产环境中的 LLM 智能体随迭代被反复评测,全量基准重跑成本高作者基于服务数万月活的 production analytics agent 的 574 次历史基准运行(按时间切分校准/留出期),比较随机采样历史缓存固定代表性子集与 IRT 自适应测试:多维 2PL 自适应测试保真度最佳只需执行 200 题(全量的 38.5%)即可将 MAE 控制在 1.03 个百分点但出于运维简单性,实际部署选择了难度分层固定子集,并验证其可免重校准迁移到另外 5 个智能体家族在短至 1 天的校准窗口内保持稳定论文以一线部署经验收尾,给出生产智能体常态化评测的实操建议

打开原文回到归档

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

中文导读

生产环境中的 LLM 智能体随迭代被反复评测,全量基准重跑成本高作者基于服务数万月活的 production analytics agent 的 574 次历史基准运行(按时间切分校准/留出期),比较随机采样历史缓存固定代表性子集与 IRT 自适应测试:多维 2PL 自适应测试保真度最佳只需执行 200 题(全量的 38.5%)即可将 MAE 控制在 1.03 个百分点但出于运维简单性,实际部署选择了难度分层固定子集,并验证其可免重校准迁移到另外 5 个智能体家族在短至 1 天的校准窗口内保持稳定论文以一线部署经验收尾,给出生产智能体常态化评测的实操建议

为什么值得关注

来自真实生产环境的一手评测经验:574 次历史运行对比四种省题策略,直接回答「生产智能体常态化评测到底要跑多少题」这一运维问题。

关键信息

  • 论文标题: Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
  • 作者: Yining She, Lei Lin
  • arXiv: https://arxiv.org/abs/2609.21267
  • 发布时间: 2026-09-18
  • arXiv 分类: cs.AI, cs.SE
  • Comments: A study of efficient recurring evaluation of a production LLM agent based on real-world historical data
  • 关联标签: agent-evaluation, production, adaptive-testing, benchmark

English Abstract

Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching, fixed representative subsets, and IRT-based adaptive testing. The results show that multidimensional 2PL adaptive testing achieves the best overall score fidelity: executing 200 questions, 38.5% of a full run, yields 1.03 pp of MAE. We nevertheless deployed difficulty-stratified fixed subsets because of their operational simplicity, and show they transfer without recalibration to five other agent families and remain stable across calibration windows as short as one day. Drawing on this deployment experience, we report practical recommendations for recurring production-agent evaluation.

English Summary

Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching, fixed representative subsets, and IRT-based adaptive testing. The results show that multidimensional 2PL adaptive testing achieves the best overall score fidelity: executing 200 questions, 38.5% of a full run, yields 1.03 pp of MAE....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。