基础设施 4.0 · 优秀 2026-09-09 · 论文

ConvMem: Convolutional Memory for Long-Context Reasoning

MemAgent 一类的顺序方法靠分段阅读+迭代更新固定大小记忆来扩展有效上下文,但延迟高且依赖昂贵的 RL 训练(容易在特定数据集上过拟合) ConvMem 把长上下文重构成层级卷积:把"带着 query 的 LLM"当作卷积核,层级式摘要文本段,把线性链式推理路径缩成对数树 三个机制:Configurable Strides 与 Skip Connections 保证证据的稳健捕获与传播;Multi-Kernel Convolution 把复杂 query 分解到解耦的语义通道 training-free可在文本段和推理线程两个维度大规模并行;RULER-HotpotQA 和 RULER-2WikiMultiHopQA 上超过 training-free 基线,并避开了 RL 训练模型在分布外任务上过拟合参数先验的风险

打开原文回到归档

ConvMem: Convolutional Memory for Long-Context Reasoning

  • ID: d34c6cc7
  • 原文链接: https://arxiv.org/abs/2609.10441
  • PDF: https://arxiv.org/pdf/2609.10441v1
  • 作者: Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu
  • 日期: 2026-09-09
  • 更新: 2026-09-09
  • 分类: infra
  • 来源类型: paper
  • 标签: long-context, memory, training-free, parallel-inference, field-note
  • 质量评分: 4/5
  • 抓取时间: 2026-09-11T04:23:41

中文导读

把长上下文推理当层级卷积:带 query 的 LLM 就是卷积核,线性链缩成对数树,training-free 且双维度并行绕开 MemAgent 式 RL 训练的延迟与过拟合

为什么值得关注

把长上下文推理当层级卷积:带 query 的 LLM 就是卷积核,线性链缩成对数树,training-free 且双维度并行绕开 MemAgent 式 RL 训练的延迟与过拟合

关键信息

  • 论文标题:ConvMem: Convolutional Memory for Long-Context Reasoning
  • 作者:Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu
  • arXiv:https://arxiv.org/abs/2609.10441
  • 发布时间:2026-09-09
  • arXiv 分类:cs.AI, cs.CL
  • 关联标签:long-context, memory, training-free, parallel-inference, field-note

English Abstract

While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory. However, this sequential paradigm suffers from high latency and requires costly reinforcement learning (RL) training, which can lead to overfitting on specific datasets. To overcome these limitations, we propose ConvMem, a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution. Inspired by CNNs, ConvMem treats an LLM prompted with a specific query as a convolutional kernel. This kernel summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Specifically, ConvMem integrates \textit{Configurable Strides} and \textit{Skip Connections} to ensure robust evidence capture and propagation, while employing \textit{Multi-Kernel Convolution} to decompose complex queries into disentangled semantic channels. This design not only mitigates error accumulation but also enables massive parallelization across both text segments and reasoning threads. Experiments on RULER-HotpotQA and RULER-2WikiMultiHopQA demonstrate that ConvMem outperforms training-free baselines and avoids the risk of overfitting to parametric priors often observed in RL-trained models on out-of-distribution tasks.

English Summary

ConvMem is a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution: an LLM prompted with a specific query acts as a convolutional kernel that summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Configurable Strides and Skip Connections ensure robust evidence capture and propagation, while Multi-Kernel Convolution decomposes complex queries into disentangled semantic channels. On RULER-HotpotQA and RULER-2WikiMultiHopQA, ConvMem outperforms training-free baselines and avoids the overfitting-to-parametric-priors risk observed in RL-trained models on out-of-distribution tasks.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。