PELM: Power-Efficient On-Device LLM Inference via Speculative Decoding + DVFS
Abstract (opencli arxiv paper)
arXiv 2609.09662 abstract (opencli arxiv paper 2609.09662 -f json): PELM jointly searches speculative decoding, variable verification depth, and DVFS frequency for on-device LLM inference. Up to 23.1% speedup and 52.4% energy reduction at unchanged task quality across hardware/datasets; code at github.com/imec-nu/PELM.
论文要点 (中文)
imec-nu(Weisi Yang、Stephen Xia)2026-09-09 提交,ACM/IEEE SenSys'26 录用。端侧 LLM 的功耗与热墙是当前最大硬约束——SoC 没风扇,高负载下要么降频要么烫手。前人 DVFS 大多只调处理器频率这一维,在热受限场景下收益有限。PELM 把「token 不需要全深度推理」拆成两个新维度:推测解码(小模型先猜几个 token、大模型整段验证)和可变验证深度(验证阶段允许提前拒绝/早停),与 DVFS 频率一起做多维搜索。端侧 LLM 推理最高加速 23.1%、能耗降低 52.4%,任务质量持平;代码在 github.com/imec-nu/PELM 开源。SenSys 体系结构/嵌入式系统会议,对做端侧 LLM 工程的人这是典型「先把现象讲清楚、再把工程方案写扎实」的稿子。落地点:多维 DVFS 比单维频率调档更稳,「不每 token 全深度验证」是会被反复重用的工程判断。
Key claims (English)
PELM (imec-nu; Weisi Yang, Stephen Xia; submitted 2026-09-09; accepted at ACM/IEEE SenSys'26) reframes power optimization for on-device LLMs as a 3-axis search over (i) speculative decoding, (ii) verification depth, and (iii) DVFS frequency, replacing single-axis DVFS that is known to saturate in thermally constrained SoCs. Across hardware and datasets it reports up to 23.1% inference speedup and 52.4% energy reduction at unchanged task quality. Code is open-sourced at github.com/imec-nu/PELM. For mobile/SoC engineering, the takeaway is the multidimensional search beats single-axis frequency tuning in the regime the paper measures, and the small/cheap drafter + variable-depth verifier combination is reusable across targets.
Obsidian 证据摘录
「PELM 把『token 不需要全深度推理』这一观察变成两个新的优化维度:推测解码(让小模型先猜几个 token,大模型整段验证)和可变验证深度(验证阶段允许提前拒绝或早停),与 DVFS 频率一起做多维搜索。跨硬件与数据集评估,端侧 LLM 推理最高加速 23.1%、能耗降 52.4%,任务质量持平;代码开源(github.com/imec-nu/PELM)。」——OpenClaw定时任务/论文流水线/2026-09-15-论文流水线.md L13-15
链接
- 论文:https://arxiv.org/abs/2609.09662
- 代码:https://github.com/imec-nu/PELM
- 论文流水线:OpenClaw定时任务/论文流水线/2026-09-15-论文流水线.md