RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention
- ID: 55246f62
- 原文链接: https://arxiv.org/abs/2608.08081
- PDF: https://arxiv.org/pdf/2608.08081v1
- 作者: Anthony. Lui, Mohamed. Elsaied, N. P. Savani
- 日期: 2026-08-08
- 更新: N/A
- 分类: infra
- 来源类型: paper
- 标签: kv-cache, quantization, moe, consumer-gpu, compression, metal
- 质量评分: 4/5
- 抓取时间: 2026-08-17T23:51:34+08:00
中文导读
三轴压缩把 26-120B MoE 塞进消费级硬件:按架构角色分派位宽的混合精度权重量化(dense 层 4-bit、路由专家 2-bit、高激活峰度的共享专家 8-bit)+ LRU 专家卸载 + 对角 SO(4) 旋转把激活各向同性化后再做 3-bit 标量 KV 量化。融合四 kernel 的 Metal 管线直接在 packed 3-bit 张量上做 attention,免物化全精度 KV——是执行模型改变而非单纯量化。16GB 内跑 26B-A4B/Qwen3-30B-A3B、32GB 内跑 Nemotron-H 120B,交互式 9-19 tok/s,困惑度近乎零退化。
为什么值得关注
120B MoE 进消费级硬件的三轴压缩:packed 3-bit 上直接做 attention,改执行模型而非只压权重。
收录理由:KV cache 压缩三部曲之一(执行模型轴),实测把 120B MoE 带进 32GB 交互式运行
Abstract
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring $O(d \log d)$ operations and 256 stored parameters per head versus $O(d^2)$ and 16{,}384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation ($Δ$PPL $\leq +0.0012$) and 100\% retrieval accuracy at 32K context.
元数据
- arXiv ID: 2608.08081
- 主分类: cs.NE
- 分类: cs.NE, cs.LG
- 评论: 9 Pages, 9 Tables, Initial submission into NeurIPS 2026
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.08081(2026-08-17)。