BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices
Source: https://arxiv.org/abs/2609.14643 · platform: arxiv · authors: Maoliang Li, Hailong Zou, Taohong Han, et al. (北京大学 + 之江实验室) · date: 2026-09-13
TL;DR(中文摘要)
手机跑 MoE 的瓶颈是数据搬运而非算力:DRAM 装不下权重,flash 分页读取碎片化,token 路由每次只换少量专家。BigMoMo 利用推测解码的多 token 验证窗口解耦「权重搬运」与「单 token 执行」:按接受率/路由影响/搬运成本剪枝推测分支与专家激活,按运行时共加载模式重排 flash 专家布局,把就绪专家与 NPU 计算批量并行。4 个 MoE 模型、5 个基准、2 个移动平台上 decode 平均比 on-demand autoregressive offload 快 4.83×,比最优推测 MoE 基线快 1.82×,支持到 30B 参数规模。
Summary (English)
BigMoMo decouples expert weight transfer from single-token execution using the multi-token verification window of speculative decoding: pruning speculative branches and expert activations, reordering flash expert layouts by runtime co-loading patterns, and batching ready experts with NPU compute. Decode is on average 4.83x faster than on-demand autoregressive offload and 1.82x faster than the best speculative MoE baseline, at MoE scales up to 30B parameters.
Abstract
Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse compute and memory overlap. BigMoMo decouples weight transfers from single-token execution by exploiting the multi-token verification window of speculative decoding: it prunes speculative branches and expert activations by acceptance rate, routing impact, and transfer cost; reorders expert layout on flash by runtime co-loading patterns; and batches ready experts with NPU computation. Across 4 MoE models, 5 benchmarks and 2 mobile platforms, decode averages 4.83x faster than on-demand autoregressive offloading and 1.82x faster than the best speculative MoE baseline, at MoE scales up to 30B parameters.
基本信息
| 项 | 值 | |------|------| | 论文 ID | 2609.14643 | | 发表 | 2026-09-13 | | 作者 | Maoliang Li, Hailong Zou, Taohong Han, et al. (北京大学 + 之江实验室) | | 备注 | 14 pages,15 figures,preprint,under review | | abs | <https://arxiv.org/abs/2609.14643> | | PDF | <https://arxiv.org/pdf/2609.14643v1> |
入库依据(同日 digest 交叉验证)
论文流水线 2026-09-16 入选:端侧 MoE runtime 最新工程样例,flash 专家重排可复现。