MoNe: Modular Neural Memory for Efficient Long Context Inference
- ID: f28f838c
- 原文链接: https://arxiv.org/abs/2608.17616
- PDF: https://arxiv.org/pdf/2608.17616
- 作者: Wonguk Cho, Kyubyung Chae, Tribhuvanesh Orekondy, Sunghyun Park, Hyoungwoo Park, Jeongho Kim, Arash Behboodi, Kyuwoong Hwang, Sungrack Yun
- 日期: 2026-08-18
- 更新: 2026-08-18
- 分类: cs.AI, cs.CL, cs.LG
- 来源类型: arxiv
- 标签: long-context, neural-memory, inference, arxiv, paper
- 质量评分: 4/5
- 抓取时间: 2026-08-20T15:44:49Z
中文导读
MoNe 是挂在任意冻结预训练 Transformer 上的轻量模块化神经记忆,无需重训即可支持长上下文推理。机制分两阶段:预处理阶段按固定大小分段读取上下文,用测试时学习的快权重记忆网络做层局部梯度更新;推理阶段只从 query token 生成 key/value,不再重读任何上下文 token。推理代价与上下文长度解耦:预处理 O(N)、查询 O(1),峰值 GPU 显存不随上下文长度增长。128K token 下相比直接上下文内学习计算量与峰值显存各降约 80%,参数开销仅 6.4%,并在 RULER 针堆取数与词提取基准上保持强势。
为什么值得关注
冻结骨干上的模块化神经记忆:O(1) 查询、峰值显存不随上下文增长,128K 下省 80% 算力
English Abstract
We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.
Obsidian 证据
- 来源 digest: 论文流水线 2026-08-20(评分 8.5)。
- 元数据与摘要经 opencli arxiv paper 核对;中文导读锚定摘要陈述的事实与数字。