Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
Original: https://arxiv.org/abs/2609.04168
Source platform: arxiv
Author: Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulik...
Original date: 2026-09-03
summary_zh
Yujie Zhang 等 9 月 3 日 arXiv (cs.DC) 投稿,IEEE TCAD 已接收把 ML 计算图映射到 heterogeneous SoC:在 Amlogic(ARM big.LITTLE + GPU)和黑芝麻(DLA + 2 DSP)上做 hierarchical 阶段内 + 阶段间算子并行,吞吐优先配置相对纯 pipeline 平均能效 +11.0%相对纯并行 +23.3%
summary_en
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture. Pa
one_liner
端侧 SoC LLM 推理的部署期图映射工具,与昨日 mzCache(运行时 KV 弹性)形成 complementary pair