Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study
- ID: d5d484a7
- 原文链接: https://arxiv.org/abs/2609.37828
- 作者: Kaikai Yuan, Rui Xi, Yu Liu
- 日期: 2026-09-29
- 更新: N/A
- 分类: infra
- 来源类型: paper
- 标签: moe, topology, congestion-control, simulation
- 质量评分: 3/5
- 抓取时间: 2026-10-01T15:57:56+00:00
中文导读
用 ASTRA-sim 加 NS-3 离散事件后端搭 32 GPU rank 受控矩阵:6 种服务器拓扑、4 种 TP/EP 划分、2 种 TP 集合通信算法、4 种网络/拥塞控制,共 768 次确定型模拟(DP/PP 固定为 1、4096-token 合成 Chakra trace、四种 MoE 配置)。开启反馈的子集里暴露通信占平均完成时间 89.9-95.8%;TP16EP2 的完成时间是 TP2EP16 的 3.68-4.35 倍;固定 rank mapping 时 Double Binary Tree 比 Ring 慢 28.3-83.2%;RoCE 上 DCQCN 比 HPCC 慢 23.8-35.7%。拓扑结论是条件性的——低 TP 度领先的拓扑到高 TP 度会反转。全是模拟、无实测,引用时记住这个边界,但作为集群规划的敏感性检查表够用。
论文信息
- arXiv ID: 2609.37828
- 提交日期: 2026-09-29
- arXiv 分类: cs.DC
- 作者: Kaikai Yuan, Rui Xi, Yu Liu
- 链接: https://arxiv.org/abs/2609.37828
Abstract(arXiv 原文)
Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, GPU--NIC connections, and the inter-node network. Using ASTRA-sim with the NS-3 discrete-event backend, we build a controlled matrix of 32 GPU ranks with data and pipeline parallelism fixed at one. Workloads are fixed-length 4096-token prefill-like synthetic Chakra traces from four MoE configurations. Experiments cover six server topologies, four TP/EP partitions, two TP collective algorithms, and four network/congestion-control modes, yielding 768 deterministic simulations. In the 144-configuration feedback-enabled subset per model, exposed communication accounts for 89.9%--95.8% of mean completion time. TP16EP2 requires 3.68--4.35x the mean completion time of TP2EP16. With fixed rank mapping, ASTRA-sim Double Binary Tree (DBT) incurs 28.3%--83.2% more time than Ring. InfiniBand-like High Precision Congestion Control (HPCC) is ~0.9% lower than HPCC over RDMA over Converged Ethernet (RoCE), whereas RoCE with Data Center Quantized Congestion Notification (DCQCN) is 23.8%--35.7% slower than RoCE HPCC. Topology effects are conditional: Topology~6 leads at low TP degrees but loses its advantage at high TP degrees, and additional GPUs or NICs help only when rank mapping balances traffic across injection paths. Under uniform 32-way sharding, the largest checkpoint-weight shard is ~48.75 GB per rank, so all configurations meet a 64 GB per-accelerator weight-residency criterion. Within the evaluated workload and simulator semantics, server topology, parallelism, collective implementation, and congestion control jointly determine exposed communication and completion time.
为什么值得关注
MoE 集群规划里「拓扑 × 划分 × 集合通信 × 拥塞控制」的联合影响第一次被系统扫描;全是模拟、无实测,当敏感性检查表用。
English Summary
A controlled simulation matrix (ASTRA-sim with NS-3, 32 GPU ranks, 768 deterministic runs) varying six server topologies, four TP/EP partitions, two TP collective algorithms and four network/congestion-control modes for MoE inference. Exposed communication accounts for 89.9-95.8% of mean completion time in the feedback-enabled subset; TP16EP2 takes 3.68-4.35x the time of TP2EP16, Double Binary Tree is 28.3-83.2% slower than Ring at fixed rank mapping, and DCQCN is 23.8-35.7% slower than HPCC on RoCE. Topology effects are conditional — leaders at low TP degrees lose their advantage at high degrees. Simulation-only: treat as a sensitivity checklist, not measurements.