REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
Source: https://arxiv.org/abs/2609.00049
Authors: Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang
Published: 2026-08-30
Categories: cs.LG, cs.AI
PDF: https://arxiv.org/pdf/2609.00049v1
Abstract
Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.
中文概要
Qian Zhang 等 2026-08-30 提交。PTQ 当前 SOTA 都是单层闭式二阶求解:为保持可解析性丢掉跨通道耦合、把输出行打包成组、再用整层冻结的 Hessian 一锅端——这一类『信息失配』正是精度损失主因。REAL-Q 反过来:不为了可解析性稀释目标函数,直接对齐端到端代理损失,每处理 128 列就做一次细粒度 Block-wise Gradient Descent 动态修正,配滑动窗口做跨层平滑过渡。在 LLaMA-3.1(8B/70B)和 Qwen3(0.6B–32B)的 W4A16 配置下,端到端 KL 散度相比全局二阶基线最高降 ~49%。W4A16 是端侧/笔记本大模型部署的主流精度点,部署链路短的团队可直接接。
一句话
REAL-Q 用 Block-wise Gradient Descent 端到端修 PTQ 信息失配,W4A16 配置 KL 散度最高降 ~49%