基础设施 4.0 · 优秀 2026-09-03 · 论文

Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs

Jinghao Wang 等 9 月 3 日 arXiv (cs.DC) 投稿并发多 agent workflow 暴露未来依赖与 serving-state 需求,同时运行在 heterogeneous GPU 池上(负载模型驻留可用资源均随时间变化)本文给出 latency-aware 调度策略,让多 agent workflow 在异构 GPU 上把 makespan 显著拉短(论文报告 -36.8%),并对 model residency 与迁移开销建模

打开原文回到归档

Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs

Original: https://arxiv.org/abs/2609.03335
Source platform: arxiv
Author: Jinghao Wang, Yifeng Zhang, Xiao Zhou, Yao Lu, Yihui Zhang, Xiaoyang Sun, Tianyu Wo, Xu Wang, Chunmi...
Original date: 2026-09-03

summary_zh

Jinghao Wang 等 9 月 3 日 arXiv (cs.DC) 投稿并发多 agent workflow 暴露未来依赖与 serving-state 需求,同时运行在 heterogeneous GPU 池上(负载模型驻留可用资源均随时间变化)本文给出 latency-aware 调度策略,让多 agent workflow 在异构 GPU 上把 makespan 显著拉短(论文报告 -36.8%),并对 model residency 与迁移开销建模

summary_en

Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, resource ordering, and placement must be selected according to the observed pool state. We present a prediction-guided runtime that uses workflow forecasts to construct and optimize a physical execution graph. Predictor estimates device-specific activation latency, peak memory, and model-loading cost, then propagates these predictions through workflow dependencies to forecast activation readiness and future model demand. Constructor builds semantics-preserving fusion and model-lifecycle alternatives, while Scheduler jointly optimizes their selection, placement, and execution

one_liner

把"agent workflow 调度"从均质 GPU 假设推到异构 GPU + 时变负载的实际生产场景,agent 平台栈基础设施论文