ConvMem: Convolutional Memory for Long-Context Reasoning
- ID: d34c6cc7
- 原文链接: https://arxiv.org/abs/2609.10441
- PDF: https://arxiv.org/pdf/2609.10441v1
- 作者: Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu
- 日期: 2026-09-09
- 更新: 2026-09-09
- 分类: infra
- 来源类型: paper
- 标签: long-context, memory, training-free, parallel-inference, field-note
- 质量评分: 4/5
- 抓取时间: 2026-09-11T04:23:41
中文导读
把长上下文推理当层级卷积:带 query 的 LLM 就是卷积核,线性链缩成对数树,training-free 且双维度并行绕开 MemAgent 式 RL 训练的延迟与过拟合
为什么值得关注
把长上下文推理当层级卷积:带 query 的 LLM 就是卷积核,线性链缩成对数树,training-free 且双维度并行绕开 MemAgent 式 RL 训练的延迟与过拟合
关键信息
- 论文标题:ConvMem: Convolutional Memory for Long-Context Reasoning
- 作者:Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu
- arXiv:https://arxiv.org/abs/2609.10441
- 发布时间:2026-09-09
- arXiv 分类:cs.AI, cs.CL
- 关联标签:long-context, memory, training-free, parallel-inference, field-note
English Abstract
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory. However, this sequential paradigm suffers from high latency and requires costly reinforcement learning (RL) training, which can lead to overfitting on specific datasets. To overcome these limitations, we propose ConvMem, a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution. Inspired by CNNs, ConvMem treats an LLM prompted with a specific query as a convolutional kernel. This kernel summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Specifically, ConvMem integrates \textit{Configurable Strides} and \textit{Skip Connections} to ensure robust evidence capture and propagation, while employing \textit{Multi-Kernel Convolution} to decompose complex queries into disentangled semantic channels. This design not only mitigates error accumulation but also enables massive parallelization across both text segments and reasoning threads. Experiments on RULER-HotpotQA and RULER-2WikiMultiHopQA demonstrate that ConvMem outperforms training-free baselines and avoids the risk of overfitting to parametric priors often observed in RL-trained models on out-of-distribution tasks.
English Summary
ConvMem is a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution: an LLM prompted with a specific query acts as a convolutional kernel that summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Configurable Strides and Skip Connections ensure robust evidence capture and propagation, while Multi-Kernel Convolution decomposes complex queries into disentangled semantic channels. On RULER-HotpotQA and RULER-2WikiMultiHopQA, ConvMem outperforms training-free baselines and avoids the overfitting-to-parametric-priors risk observed in RL-trained models on out-of-distribution tasks.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。