基础设施 4.0 · 优秀 2026-08-11 · 论文

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantizat...

ReRound 解决 calibration-free LLM 量化里的 midpoint ambiguity:训练 conditional diffusion model 生成 low-bit weight 的连续 reconstruction,为接近量化区间中点的 weight 提供 rounding direction guidance;非敏感权重继续 RTN方法离线推理无额外开销,小模型 3-bit / 4-bit 量化稳定超过 RTN

打开原文回到归档

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

Source: https://arxiv.org/abs/2608.11045
Author: He-Yen Hsieh, H. T. Kung
Original date: 2026-08-11
Added by: AAIF daily-intake-evening 2026-08-13

摘要

ReRound 解决 calibration-free LLM 量化里的 midpoint ambiguity:训练 conditional diffusion model 生成 low-bit weight 的连续 reconstruction,为接近量化区间中点的 weight 提供 rounding direction guidance;非敏感权重继续 RTN。方法离线、推理无额外开销,小模型 3-bit / 4-bit 量化稳定超过 RTN。

English Summary

ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.

入库理由

  • quality_score: 4
  • category: infra
  • tags: llm-quantization, on-device-ai, diffusion, inference-optimization
  • one_liner: ReRound 用 diffusion-guided rounding 改善小 LLM 的 calibration-free 低比特量化。

Obsidian evidence excerpt

eline/2026-08-13/`

## 今日论文速报

今天 arXiv recent 覆盖 `cs.AI / cs.CL / cs.LG / cs.CV / cs.RO / cs.MA` 六个分类共 254 篇去重提交。本轮最有信号的方向是 **Agent 记忆与自演化基础设施**——多条线从不同角度在做同一件事:把"retrieve-only"换成"compile + skill + provenance"。`SkillZip / Muscle Memory / GeoForge / MAP-Graph / EvoMem` 形成今天的"memory 范式切换"主线;`Self-Evolving GUI Grounding` + `SPIEval` 把这条主线拉到 GUI/移动端;`ReRound / Gated VLA-Cache` 提供 on-device 推理的量化与缓存路径。安全侧一条新的攻击面:`Trajectory Backdoor Attack` 直接攻击 self-evolving skill 的可信演化管道。

1. **SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure** — arXiv:2608.11079。把 self-evolving agent 累积的 skill 当成"typed contract"——名字/描述/工作流/工具契约/输出字段/例外规则——用 typed minimum description-length 目标一次性压缩:重复规则 state once at scope,重复动作序列 factor into shared procedure,例外保留为 explicit exceptions。Zip-on-Write 模式支持 incremental 演化不重放任务。压缩率高、保留 unique rare rules by construction。
   来源:https://arxiv.org/abs/2608.11079

2. **Muscle Memory for Agents: Compile not Merely Retrieve** — arXiv:2608.08995。主张把"反复出现的用户意图"编译成 purpose-built specialist agent,而不是 retrieve-then-orchestrate。Harvest→Analyze→Augment→Evaluate 四阶段管线,从对话历史中分别挖出 behavioral pattern 和 task pattern,发出的 specialist 配 two-stage trigger matching。90 held-out scenario 上 specialist 触发时 88.9% 胜率,+2.05 personalisation gain,accuracy 损失仅 -0.28(1-4 scale)。
   来源:https://arxiv.org/abs/2608.08995

3. **GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning** — arXiv:2608.10494。Training-free self-evolving 框架,把完成的轨迹压成结构化 nonparametric execution state——Workflow Graph Memory(全局操作顺序)+ Action-Level Experiences(局部纠错)+ Adapted Skill SOP(程序与数据约束)。执行、蒸馏、复用三段循环里 backbone LLM 不变。多个 geospatial benchmark 上同时拉高 task accuracy 和 tool-use trajectory quality。
   来源:https://arxiv.org/abs/2608.10494

4. **MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows** — arXiv:2608.10509。把 agent、source、memory、claim、action 全部建模成 typed execution graph:lineage tracing + permission-ineligible record exclusion + semantic similarity × multiplicative path trust reranking + risk-sensitive action gate。2,700 合成任务上 94.96% task success / 72.70% exact decision accuracy / 9

arXiv metadata / abstract

ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.