基础设施 4.0 · 优秀 2026-08-15 · 论文

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

AgentSysBench 用十个代表性 agentic 应用和统一的系统层埋点,在受控部署与生产 trace 上量出 agentic 负载与常规推理服务的六条差异:单会话沙箱工作集内存峰值 28 GB;十个应用里五个的非 LLM 组件占延迟大头;GPU 推理内存密集检索与 CPU 密集沙箱混部使任务延迟最多相差 32 倍;生产会话在活跃步骤之间把状态闲置数分钟到数小时;工具 schema 与观测内容构成的 control-plane 开销持续挤占有效上下文;搜索查询与网络抓取存在大量跨请求重复四项优化实验验证这些数字能兑现:任务感知服务延迟降低 29%-40%,通信感知放置最高 4.5 倍,状态卸载节省 4.6 倍内存,工具结果缓存消除 35.2% 重复搜索并节省 19.3% 搜索成本

打开原文回到归档

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
  • AgentSysBench 用十个代表性 agentic 应用和统一的系统层埋点,在受控部署与生产 trace 上量出 agentic 负载与常规推理服务的六条差异:单会话沙箱工作集内存峰值 28 GB;十个应用里五个的非 LLM 组件占延迟大头;GPU 推理、内存密集检索与 CPU 密集沙箱混部使任务延迟最多相差 32 倍;生产会话在活跃步骤之间把状态闲置数分钟到数小时;工具 schema 与观测内容构成的 control-plane 开销持续挤占有效上下文;搜索查询与网络抓取存在大量跨请求重复。四项优化实验验证这些数字能兑现:任务感知服务延迟降低 29%-40%,通信感知放置最高 4.5 倍,状态卸载节省 4.6 倍内存,工具结果缓存消除 35.2% 重复搜索并节省 19.3% 搜索成本。

论文信息

  • arXiv ID: 2608.15127
  • 作者: Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
  • 发表: 2026-08-15(更新:2026-08-15)
  • 分类: cs.OS, cs.AI, cs.DC, cs.MA
  • 原文链接: https://arxiv.org/abs/2608.15127
  • PDF: https://arxiv.org/pdf/2608.15127
  • 标签: agentic-workloads serving-systems benchmark profiling resource-management

Abstract(原文)

Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.

核心要点(英文摘要的中文提炼)

  • 论文题为 From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems,发表于 arXiv(cs.OS, cs.AI, cs.DC, cs.MA,2026-08-15 提交/更新)。
  • 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。

Obsidian 证据摘录

入选自 Obsidian《论文流水线 · 2026-08-31》速报第2篇并列为值得精读:agentic 负载系统画像,生产 trace,数字全。