模型与实验室 4.0 · 优秀 2026-08-21 · 论文

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

CoT 推理的长思考链是推理成本大头,压缩太狠又会破坏逻辑连贯这篇提出上下文-生成替代律:把历史推理轨迹中可复用的推理模式关键约束关键操作提炼成记忆,作为 prefill 侧脚手架在压缩后补回信息免训练框架,在 GSM8K/MATH/BBH/MMLU-Sci 上比 Chain-of-Draft 压缩基线分别提 21.4/28.0/29.5/6.61 个点,同时相对标准 CoT 有 1.14-1.49 合的延迟加速,且与各级压缩机制兼容对推理成本敏感的 agent 部署是可行配方

打开原文回到归档

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

  • ID: f326068a
  • 原文链接: https://arxiv.org/abs/2608.21265
  • PDF: https://arxiv.org/pdf/2608.21265v1
  • 作者: Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang, Tingwen Liu
  • 日期: 2026-08-21
  • 更新: 2026-08-21
  • 分类: models
  • 来源类型: paper
  • 标签: cot, inference-efficiency, agent-memory, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-25T15:40:00+00:00

中文导读

CoT 推理的长思考链是推理成本大头,压缩太狠又会破坏逻辑连贯。这篇提出「上下文-生成替代律」:把历史推理轨迹中可复用的推理模式、关键约束、关键操作提炼成记忆,作为 prefill 侧脚手架在压缩后补回信息。免训练框架,在 GSM8K/MATH/BBH/MMLU-Sci 上比 Chain-of-Draft 压缩基线分别提 21.4/28.0/29.5/6.61 个点,同时相对标准 CoT 有 1.14-1.49 合的延迟加速,且与各级压缩机制兼容。对推理成本敏感的 agent 部署是可行配方。

为什么值得关注

CoT 推理的长思考链是推理成本大头,压缩太狠又会破坏逻辑连贯。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning
  • 作者:Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang, Tingwen Liu
  • arXiv:https://arxiv.org/abs/2608.21265
  • 发布时间:2026-08-21
  • arXiv 分类:cs.CL

English Abstract

Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may disrupt logical coherence and degrade performance. We formalize this trade-off as the \textit{Context-Generation Substitution Law}, where explicit reasoning context substitutes for part of decode-time generation. Based on this principle, we propose \textit{Memory-Augmented Compression}, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds. Rather than using raw demonstrations, these memories summarize reusable reasoning patterns, key constraints, and critical operations to compensate for information lost during compression. Experiments show that Memory consistently improves prompt-based Chain-of-Draft (CoD) compression across mathematical reasoning, complex reasoning, and science question answering tasks, yielding accuracy gains of 21.4, 28.0, 29.5, and 6.61 points over CoD on GSM8K, MATH, BBH, and MMLU-Sci, while achieving a 1.14--1.49$\times$ latency speedup over standard CoT. Memory is also compatible with token-level, reasoning-trace-level, and inference-state compression mechanisms. Further analyzes show that the gains come from relevant reasoning memories rather than simply increasing context length.

English Summary

Formalizes a context-generation substitution law for CoT: distill reusable reasoning patterns, key constraints and operations from historical traces into memory, re-injected as prefill-side scaffolding after compression. Training-free, +21.4/28.0/29.5/6.61 points over Chain-of-Draft on GSM8K/MATH/BBH/MMLU-Sci with 1.14-1.49x latency speedup vs standard CoT.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。