Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
- ID: f931afb5
- 原文链接: https://arxiv.org/abs/2608.20953
- PDF: https://arxiv.org/pdf/2608.20953v1
- 作者: Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
- 日期: 2026-08-21
- 更新: 2026-08-21
- 分类: models
- 来源类型: paper
- 标签: quantization, distillation, inference-efficiency, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-25T15:40:00+00:00
中文导读
结构压缩 + 4-bit 量化的模型在推理、数学、代码、长上下文上掉点严重,默认恢复手段 QAT 收敛慢且过峰即崩。QAH 换了蒸馏对象:压缩模型的 bfloat16 检查点本来就是对原模型的蒸馏近似,不如让 4-bit 学生直接从原始未压缩模型蒸馏。在 GPT-OSS 120B 压到 60B 再压到 MXFP4 的管线上,QAH 学生在 9 个基准中的 7 个上追平或超过其 bfloat16 源,内存约为四分之一,以开权重发布为 Hypernova-60B;对比 QAT 基线约 7 倍速度达到可比峰值且继续训练稳定。一份不需要数周超参搜索就能落地的压缩恢复配方。
为什么值得关注
结构压缩 + 4-bit 量化的模型在推理、数学、代码、长上下文上掉点严重,默认恢复手段 QAT 收敛慢且过峰即崩。 实验与数字均来自论文摘要本身。
关键信息
- 论文标题:Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
- 作者:Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
- arXiv:https://arxiv.org/abs/2608.20953
- 发布时间:2026-08-21
- arXiv 分类:cs.CL, cs.AI, cs.LG, cs.PF
English Abstract
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.
English Summary
Instead of healing a compressed+4-bit model from its own bfloat16 checkpoint, QAH distills the 4-bit student directly from the original uncompressed model. On the GPT-OSS 120B to 60B to MXFP4 pipeline, the student matches or beats its bfloat16 source on 7 of 9 benchmarks at ~1/4 weight memory (released as Hypernova-60B), reaching comparable peaks ~7x faster than matched QAT baselines and remaining stable under continued training.
Obsidian Notes
- 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
- 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。