AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference
Original: https://arxiv.org/abs/2609.02514
Source platform: arxiv
Author: Yida Zhang, Zhiyong Gao, Shuaibing Yue, Jie Li, Rui Wang
Original date: 2026-09-02
summary_zh
Yida Zhang 等 9 月 2 日 arXiv (cs.DC) 投稿端侧部署 LLM 通常依赖模型压缩或 split inference,但压缩会损推理能力,split inference 受网络不稳拖累AceSpec 提出非对称端云协同框架,WAN 不稳时把 speculative pipeline flush 转本地 O(1) 查询;3.52 加速,50 Kbps 极限场景近峰值吞吐
summary_en
Deploying Large Language Models (LLMs) on edge devices typically relies on model compression or split inference. However, compression degrades reasoning capabilities, while split inference suffers from severe Wide Area Network (WAN) communication bottlenecks. Edge-cloud speculative decoding emerges as a promising alternative, leveraging an edge small model to draft tokens for cloud verification. Yet, over volatile WANs, inevitable prediction rejections trigger catastrophic pipeline stalls and network-wide rollbacks, neutralizing collaborative gains. To overcome this, we propose AceSpec, an asymmetric edge-cloud collaborative framework. AceSpec utilizes un-saturated edge compute to proactively construct a probabilistic state cache, effectively transforming network-wide pipeline flushes into $\mathcal{O}(1)$ local memory lookups. To preserve bandwidth, it employs an asymmetric communicatio
one_liner
WAN 不稳下 speculative decoding 也能跑:pipeline flush 转本地 O(1) 查询,端云协同解码的实用锚点