模型与实验室 3.0 · 值得看 2026-09-29 · 论文

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

KV 量化的 2-bit 区间里变换选择决定误差WUSH-KV 把 WUSH 变换用于 KV:用校准数据从矩阵乘两个因子的二阶统计构造数据感知变换,key/value 分开处理value 变换折进模型权重key 变换在 RoPE 后施加,可搭配 clipped 量化器;对 QuEST INT 量化器证明了 WUSH 变换在温和假设下近似最优端到端集成进 SGLang(OSCAR 式 percentile-clipped affine 量化),2-bit 下在全部受测模型与下游任务上与 OSCAR 变换持平或更好,层内重建误差与端到端困惑度最低

打开原文回到归档

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

  • ID: f375b2ed
  • 原文链接: https://arxiv.org/abs/2609.38121
  • 作者: Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh
  • 日期: 2026-09-29
  • 更新: N/A
  • 分类: models
  • 来源类型: paper
  • 标签: kv-cache, quantization, long-context, sglang
  • 质量评分: 3/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

KV 量化的 2-bit 区间里变换选择决定误差。WUSH-KV 把 WUSH 变换用于 KV:用校准数据从矩阵乘两个因子的二阶统计构造数据感知变换,key/value 分开处理——value 变换折进模型权重、key 变换在 RoPE 后施加,可搭配 clipped 量化器;对 QuEST INT 量化器证明了 WUSH 变换在温和假设下近似最优。端到端集成进 SGLang(OSCAR 式 percentile-clipped affine 量化),2-bit 下在全部受测模型与下游任务上与 OSCAR 变换持平或更好,层内重建误差与端到端困惑度最低。

论文信息

  • arXiv ID: 2609.38121
  • 提交日期: 2026-09-29
  • arXiv 分类: cs.LG
  • 作者: Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh
  • 链接: https://arxiv.org/abs/2609.38121

Abstract(arXiv 原文)

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.

为什么值得关注

KV 压缩这条线里,驱逐(ActKV)与位宽(本篇)是正交的两刀、可叠加;2-bit 区间「变换选择」的近优性第一次有了证明。

English Summary

WUSH-KV applies the WUSH data-aware transform — constructed from second-order statistics of both factors in a matrix product — to low-bit KV quantization, with separate key/value transforms (value folded into model weights, key applied after RoPE) that pair with clipped quantizers; for QuEST INT the WUSH transform is near-optimal under mild assumptions. Integrated into SGLang with OSCAR-style percentile-clipped affine quantization, 2-bit WUSH-KV matches or beats the OSCAR transform across all evaluated models and downstream tasks, with the lowest layerwise reconstruction error and end-to-end perplexity.