基础设施 4.0 · 优秀 2026-08-08 · 论文

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

三轴压缩把 26-120B MoE 塞进消费级硬件:按架构角色分派位宽的混合精度权重量化(dense 层 4-bit路由专家 2-bit高激活峰度的共享专家 8-bit)+ LRU 专家卸载 + 对角 SO(4) 旋转把激活各向同性化后再做 3-bit 标量 KV 量化融合四 kernel 的 Metal 管线直接在 packed 3-bit 张量上做 attention,免物化全精度 KV是执行模型改变而非单纯量化16GB 内跑 26B-A4B/Qwen3-30B-A3B32GB 内跑 Nemotron-H 120B,交互式 9-19 tok/s,困惑度近乎零退化

打开原文回到归档

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

  • ID: 55246f62
  • 原文链接: https://arxiv.org/abs/2608.08081
  • PDF: https://arxiv.org/pdf/2608.08081v1
  • 作者: Anthony. Lui, Mohamed. Elsaied, N. P. Savani
  • 日期: 2026-08-08
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: kv-cache, quantization, moe, consumer-gpu, compression, metal
  • 质量评分: 4/5
  • 抓取时间: 2026-08-17T23:51:34+08:00

中文导读

三轴压缩把 26-120B MoE 塞进消费级硬件:按架构角色分派位宽的混合精度权重量化(dense 层 4-bit、路由专家 2-bit、高激活峰度的共享专家 8-bit)+ LRU 专家卸载 + 对角 SO(4) 旋转把激活各向同性化后再做 3-bit 标量 KV 量化。融合四 kernel 的 Metal 管线直接在 packed 3-bit 张量上做 attention,免物化全精度 KV——是执行模型改变而非单纯量化。16GB 内跑 26B-A4B/Qwen3-30B-A3B、32GB 内跑 Nemotron-H 120B,交互式 9-19 tok/s,困惑度近乎零退化。

为什么值得关注

120B MoE 进消费级硬件的三轴压缩:packed 3-bit 上直接做 attention,改执行模型而非只压权重。

收录理由:KV cache 压缩三部曲之一(执行模型轴),实测把 120B MoE 带进 32GB 交互式运行

Abstract

Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring $O(d \log d)$ operations and 256 stored parameters per head versus $O(d^2)$ and 16{,}384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation ($Δ$PPL $\leq +0.0012$) and 100\% retrieval accuracy at 32K context.

元数据

  • arXiv ID: 2608.08081
  • 主分类: cs.NE
  • 分类: cs.NE, cs.LG
  • 评论: 9 Pages, 9 Tables, Initial submission into NeurIPS 2026
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.08081(2026-08-17)。