Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution
Source: https://arxiv.org/abs/2609.14213 · platform: arxiv · authors: Zhuojin Li, Marco Paolieri, Leana Golubchik · date: 2026-09-13
TL;DR(中文摘要)
移动端推理跑在多核 CPU 集群 + 移动 GPU 的异构平台上,现有优化要么算子间并行(整个算子分给 CPU 或 GPU),要么算子内并行(切分每个算子),没有统一答案。本文把 partition、device assignment、execution order 三个决策联合建模求解,给出 partition-aware 调度器,在真机上端到端延迟优于两类纯策略与 SOTA 基线。IFIP Performance 2026 / Performance Evaluation 期刊录用。对端侧 LLM 推理调度是 CPU-GPU DAG 协同的参考实现。
Summary (English)
Jointly optimizing partition, device assignment, and execution order for latency-critical DNN inference on mobile CPU-cluster + GPU platforms beats both pure inter-operator and pure intra-operator parallelism baselines. Accepted at IFIP Performance 2026 (Performance Evaluation journal).
Abstract
Modern mobile inference runs on heterogeneous platforms combining mobile GPUs with multiple CPU core clusters. Existing optimizations typically exploit either inter-operator parallelism, by assigning entire operators to CPU cores or to the GPU, or intra-operator parallelism, by partitioning each operator. This paper studies the joint problem of partition, device assignment, and execution order for latency-critical DNN inference on mobile heterogeneous platforms, and shows that neither pure inter- nor pure intra-operator strategies are uniformly optimal. The proposed partition-aware scheduler co-optimizes the three decisions, improving end-to-end latency over state-of-the-art baselines on real devices.
基本信息
| 项 | 值 | |------|------| | 论文 ID | 2609.14213 | | 发表 | 2026-09-13 | | 作者 | Zhuojin Li, Marco Paolieri, Leana Golubchik | | 备注 | Accepted for publication in the Performance Evaluation journal (presented at IFIP Performance 2026, in Ghent, Belgium) | | abs | <https://arxiv.org/abs/2609.14213> | | PDF | <https://arxiv.org/pdf/2609.14213v1> |
入库依据(同日 digest 交叉验证)
论文流水线 2026-09-16 备选中等相关:移动异构 CPU-GPU DAG 调度;opencli 元数据核实。