模型与实验室 3.0 · 值得看 2026-09-29 · 论文

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

GDNKDA 这类线性注意力把上下文压进固定大小循环状态,状态的反复读写成为新推理瓶颈;直接低比特量化又因舍入误差累积与 outlier 行列崩精度LeapQuant 免训练做到 8-bit 近乎无损:按窗口跳量窗口内用低比特状态加高精度缓冲更新算输出,抑制误差累积;保留状态中最大的 outlier 成少数个高精度 Compensator Token(与真实 token 共享更新路径),再对残差平滑后量化QwenKimiGLM 三个模型家族上精度对齐 FP32 基线,内核级加速 2.05-3.70 倍,B200/RTX PRO 6000/RTX 5090 端到端 1.47 倍

打开原文回到归档

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

  • ID: 4b189b24
  • 原文链接: https://arxiv.org/abs/2609.38166
  • 作者: Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
  • 日期: 2026-09-29
  • 更新: N/A
  • 分类: models
  • 来源类型: paper
  • 标签: linear-attention, quantization, recurrent-state, inference
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

GDN、KDA 这类线性注意力把上下文压进固定大小循环状态,状态的反复读写成为新推理瓶颈;直接低比特量化又因舍入误差累积与 outlier 行列崩精度。LeapQuant 免训练做到 8-bit 近乎无损:按窗口跳量、窗口内用低比特状态加高精度缓冲更新算输出,抑制误差累积;保留状态中最大的 outlier 成少数个高精度 Compensator Token(与真实 token 共享更新路径),再对残差平滑后量化。Qwen、Kimi、GLM 三个模型家族上精度对齐 FP32 基线,内核级加速 2.05-3.70 倍,B200/RTX PRO 6000/RTX 5090 端到端 1.47 倍。

论文信息

  • arXiv ID: 2609.38166
  • 提交日期: 2026-09-29
  • arXiv 分类: cs.LG, cs.AI
  • 作者: Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
  • 链接: https://arxiv.org/abs/2609.38166

Abstract(arXiv 原文)

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.

为什么值得关注

线性注意力铺开之后,循环状态的量化会跟 KV 量化一样成为标配工程问题;免训练加三大家族验证是落地友好信号。

English Summary

LeapQuant is a training-free scheme achieving near-lossless 8-bit quantization of the fixed-size recurrent states in linear-attention hybrids (GDN, KDA): per-window leap quantization computes outputs from a low-bit state plus high-precision buffered updates to limit error accumulation, while the state's largest outliers are retained as high-precision Compensator Tokens sharing the real-token update path, with residual smoothing before quantization. Across Qwen, Kimi and GLM families it matches FP32 accuracy with 2.05-3.70x kernel-level and 1.47x end-to-end speedups on B200, RTX PRO 6000 and RTX 5090.