模型与实验室 5.0 · 必读 2026-09-09 · 论文

PELM: Power-Efficient On-Device LLM Inference via Speculative Decoding + DVFS

PELM(imec-nu, 2026-09-09 提交,ACM/IEEE SenSys'26 录用)针对端侧 LLM 的功耗与热墙问题,把token 不需要每步都做全深度推理这一观察拆成两个新维度:推测解码(小模型先猜几个 token,大模型整段验证)和可变验证深度(验证阶段允许提前拒绝或早停),与 DVFS 频率一起做多维搜索论文跨硬件/数据集评估,端侧 LLM 推理最高加速 23.1%能耗降低 52.4%,任务质量持平;代码在 github.com/imec-nu/PELM对做端侧 LLM 工程的人来说,多维 DVFS 比单维频率调档更稳,不每 token 都全深度验证是会反复重用的工程判断

打开原文回到归档

PELM: Power-Efficient On-Device LLM Inference via Speculative Decoding + DVFS

Abstract (opencli arxiv paper)

arXiv 2609.09662 abstract (opencli arxiv paper 2609.09662 -f json): PELM jointly searches speculative decoding, variable verification depth, and DVFS frequency for on-device LLM inference. Up to 23.1% speedup and 52.4% energy reduction at unchanged task quality across hardware/datasets; code at github.com/imec-nu/PELM.

论文要点 (中文)

imec-nu(Weisi Yang、Stephen Xia)2026-09-09 提交,ACM/IEEE SenSys'26 录用。端侧 LLM 的功耗与热墙是当前最大硬约束——SoC 没风扇,高负载下要么降频要么烫手。前人 DVFS 大多只调处理器频率这一维,在热受限场景下收益有限。PELM 把「token 不需要全深度推理」拆成两个新维度:推测解码(小模型先猜几个 token、大模型整段验证)和可变验证深度(验证阶段允许提前拒绝/早停),与 DVFS 频率一起做多维搜索。端侧 LLM 推理最高加速 23.1%、能耗降低 52.4%,任务质量持平;代码在 github.com/imec-nu/PELM 开源。SenSys 体系结构/嵌入式系统会议,对做端侧 LLM 工程的人这是典型「先把现象讲清楚、再把工程方案写扎实」的稿子。落地点:多维 DVFS 比单维频率调档更稳,「不每 token 全深度验证」是会被反复重用的工程判断。

Key claims (English)

PELM (imec-nu; Weisi Yang, Stephen Xia; submitted 2026-09-09; accepted at ACM/IEEE SenSys'26) reframes power optimization for on-device LLMs as a 3-axis search over (i) speculative decoding, (ii) verification depth, and (iii) DVFS frequency, replacing single-axis DVFS that is known to saturate in thermally constrained SoCs. Across hardware and datasets it reports up to 23.1% inference speedup and 52.4% energy reduction at unchanged task quality. Code is open-sourced at github.com/imec-nu/PELM. For mobile/SoC engineering, the takeaway is the multidimensional search beats single-axis frequency tuning in the regime the paper measures, and the small/cheap drafter + variable-depth verifier combination is reusable across targets.

Obsidian 证据摘录

「PELM 把『token 不需要全深度推理』这一观察变成两个新的优化维度:推测解码(让小模型先猜几个 token,大模型整段验证)和可变验证深度(验证阶段允许提前拒绝或早停),与 DVFS 频率一起做多维搜索。跨硬件与数据集评估,端侧 LLM 推理最高加速 23.1%、能耗降 52.4%,任务质量持平;代码开源(github.com/imec-nu/PELM)。」——OpenClaw定时任务/论文流水线/2026-09-15-论文流水线.md L13-15

链接