基础设施 3.0 · 值得看 2026-09-28 · 论文

Semantics, Workflows, and Infrastructure: Understanding Agent Serving at Production Scale

分析某通用 agent 生产平台两周 1170 万请求的 trace底层超 1 万 GPU 的推理集群,把测量接成三层:任务层发生语义工作流层执行形态基础设施层 serving 需求发现的负载形态包括:会话间请求量高度倾斜逻辑兄弟请求很少出现执行重叠任务边界之间存在上下文复用这些数字是 agent serving 调度与容量设计的直接输入,也让 EfficientAgent 的 working set 结论有了生产侧印证;表征研究本身不提供方案,但 open problems 列得比较具体

打开原文回到归档

Semantics, Workflows, and Infrastructure: Understanding Agent Serving at Production Scale

  • ID: 2c1ab38a
  • 原文链接: https://arxiv.org/abs/2609.34432
  • 作者: Yihao Zheng, Jingzhe Jiang, Dejiang Zhu, Zhiyuan Tan, Yang Tian, Tao Wang, Minchen Yu
  • 日期: 2026-09-28
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: agent-serving, characterization, production-trace, capacity-planning
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

分析某通用 agent 生产平台两周 1170 万请求的 trace、底层超 1 万 GPU 的推理集群,把测量接成三层:任务层发生语义、工作流层执行形态、基础设施层 serving 需求。发现的负载形态包括:会话间请求量高度倾斜、逻辑兄弟请求很少出现执行重叠、任务边界之间存在上下文复用。这些数字是 agent serving 调度与容量设计的直接输入,也让 EfficientAgent 的 working set 结论有了生产侧印证;表征研究本身不提供方案,但 open problems 列得比较具体。

论文信息

  • arXiv ID: 2609.34432
  • 提交日期: 2026-09-28
  • arXiv 分类: cs.DC
  • 作者: Yihao Zheng, Jingzhe Jiang, Dejiang Zhu, Zhiyuan Tan, Yang Tian, Tao Wang, Minchen Yu
  • 链接: https://arxiv.org/abs/2609.34432

Abstract(arXiv 原文)

Large language model (LLM) agents execute applications through a workflow of inference requests with tool calls and user interactions. Serving these applications at production scale requires understanding how application behavior shapes inference demand and for guiding efficient execution. Recent characterization studies provide request-level workload measurements and agent execution analysis. However, an end-to-end view connecting task initiation, workflow execution, and inference infrastructure remains unexplored. In this paper, we analyze a two-week trace of 11.7 million requests from a large-scale production platform for general-purpose agents, backed by inference infrastructure comprising over 10k GPUs. We characterize the platform at three connected levels: task-level initiation semantics, workflow-level execution patterns, and infrastructure level serving demands. Our measurements reveal workload patterns such as highly skewed request volumes across sessions, rare execution overlap among logical sibling requests, and context reuse across task boundaries. Building on these observations, we analyze deployment implications and identify open problems to guide future research on agent serving systems.

为什么值得关注

生产侧真实负载形态是调度与容量设计的直接输入;把这篇与 EfficientAgent 的 working set 判据对读,能对出「实验室机制 vs 生产形态」的差距。

English Summary

A characterization of a two-week, 11.7M-request trace from a production general-purpose agent platform backed by 10k+ inference GPUs, connecting task-level initiation semantics, workflow-level execution patterns, and infrastructure-level serving demands. Measured patterns include highly skewed request volumes across sessions, rare execution overlap among logical sibling requests, and context reuse across task boundaries, with deployment implications and concrete open problems for agent serving systems.