模型与实验室 4.0 · 优秀 2026-08-21 · 论文

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

结构压缩 + 4-bit 量化的模型在推理数学代码长上下文上掉点严重,默认恢复手段 QAT 收敛慢且过峰即崩QAH 换了蒸馏对象:压缩模型的 bfloat16 检查点本来就是对原模型的蒸馏近似,不如让 4-bit 学生直接从原始未压缩模型蒸馏在 GPT-OSS 120B 压到 60B 再压到 MXFP4 的管线上,QAH 学生在 9 个基准中的 7 个上追平或超过其 bfloat16 源,内存约为四分之一,以开权重发布为 Hypernova-60B;对比 QAT 基线约 7 倍速度达到可比峰值且继续训练稳定一份不需要数周超参搜索就能落地的压缩恢复配方

打开原文回到归档

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

  • ID: f931afb5
  • 原文链接: https://arxiv.org/abs/2608.20953
  • PDF: https://arxiv.org/pdf/2608.20953v1
  • 作者: Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
  • 日期: 2026-08-21
  • 更新: 2026-08-21
  • 分类: models
  • 来源类型: paper
  • 标签: quantization, distillation, inference-efficiency, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-25T15:40:00+00:00

中文导读

结构压缩 + 4-bit 量化的模型在推理、数学、代码、长上下文上掉点严重,默认恢复手段 QAT 收敛慢且过峰即崩。QAH 换了蒸馏对象:压缩模型的 bfloat16 检查点本来就是对原模型的蒸馏近似,不如让 4-bit 学生直接从原始未压缩模型蒸馏。在 GPT-OSS 120B 压到 60B 再压到 MXFP4 的管线上,QAH 学生在 9 个基准中的 7 个上追平或超过其 bfloat16 源,内存约为四分之一,以开权重发布为 Hypernova-60B;对比 QAT 基线约 7 倍速度达到可比峰值且继续训练稳定。一份不需要数周超参搜索就能落地的压缩恢复配方。

为什么值得关注

结构压缩 + 4-bit 量化的模型在推理、数学、代码、长上下文上掉点严重,默认恢复手段 QAT 收敛慢且过峰即崩。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
  • 作者:Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
  • arXiv:https://arxiv.org/abs/2608.20953
  • 发布时间:2026-08-21
  • arXiv 分类:cs.CL, cs.AI, cs.LG, cs.PF

English Abstract

Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.

English Summary

Instead of healing a compressed+4-bit model from its own bfloat16 checkpoint, QAH distills the 4-bit student directly from the original uncompressed model. On the GPT-OSS 120B to 60B to MXFP4 pipeline, the student matches or beats its bfloat16 source on 7 of 9 benchmarks at ~1/4 weight memory (released as Hypernova-60B), reaching comparable peaks ~7x faster than matched QAT baselines and remaining stable under continued training.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。