From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- AgentSysBench 用十个代表性 agentic 应用和统一的系统层埋点,在受控部署与生产 trace 上量出 agentic 负载与常规推理服务的六条差异:单会话沙箱工作集内存峰值 28 GB;十个应用里五个的非 LLM 组件占延迟大头;GPU 推理、内存密集检索与 CPU 密集沙箱混部使任务延迟最多相差 32 倍;生产会话在活跃步骤之间把状态闲置数分钟到数小时;工具 schema 与观测内容构成的 control-plane 开销持续挤占有效上下文;搜索查询与网络抓取存在大量跨请求重复。四项优化实验验证这些数字能兑现:任务感知服务延迟降低 29%-40%,通信感知放置最高 4.5 倍,状态卸载节省 4.6 倍内存,工具结果缓存消除 35.2% 重复搜索并节省 19.3% 搜索成本。
论文信息
- arXiv ID: 2608.15127
- 作者: Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
- 发表: 2026-08-15(更新:2026-08-15)
- 分类: cs.OS, cs.AI, cs.DC, cs.MA
- 原文链接: https://arxiv.org/abs/2608.15127
- PDF: https://arxiv.org/pdf/2608.15127
- 标签:
agentic-workloadsserving-systemsbenchmarkprofilingresource-management
Abstract(原文)
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.
核心要点(英文摘要的中文提炼)
- 论文题为 From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems,发表于 arXiv(cs.OS, cs.AI, cs.DC, cs.MA,2026-08-15 提交/更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
Obsidian 证据摘录
入选自 Obsidian《论文流水线 · 2026-08-31》速报第2篇并列为值得精读:agentic 负载系统画像,生产 trace,数字全。