模型与实验室 4.0 · 优秀 2026-08-18 · 论文

MoNe: Modular Neural Memory for Efficient Long Context Inference

MoNe 是挂在任意冻结预训练 Transformer 上的轻量模块化神经记忆,无需重训即可支持长上下文推理机制分两阶段:预处理阶段按固定大小分段读取上下文,用测试时学习的快权重记忆网络做层局部梯度更新;推理阶段只从 query token 生成 key/value,不再重读任何上下文 token推理代价与上下文长度解耦:预处理 O(N)查询 O(1),峰值 GPU 显存不随上下文长度增长128K token 下相比直接上下文内学习计算量与峰值显存各降约 80%,参数开销仅 6.4%,并在 RULER 针堆取数与词提取基准上保持强势

打开原文回到归档

MoNe: Modular Neural Memory for Efficient Long Context Inference

  • ID: f28f838c
  • 原文链接: https://arxiv.org/abs/2608.17616
  • PDF: https://arxiv.org/pdf/2608.17616
  • 作者: Wonguk Cho, Kyubyung Chae, Tribhuvanesh Orekondy, Sunghyun Park, Hyoungwoo Park, Jeongho Kim, Arash Behboodi, Kyuwoong Hwang, Sungrack Yun
  • 日期: 2026-08-18
  • 更新: 2026-08-18
  • 分类: cs.AI, cs.CL, cs.LG
  • 来源类型: arxiv
  • 标签: long-context, neural-memory, inference, arxiv, paper
  • 质量评分: 4/5
  • 抓取时间: 2026-08-20T15:44:49Z

中文导读

MoNe 是挂在任意冻结预训练 Transformer 上的轻量模块化神经记忆,无需重训即可支持长上下文推理。机制分两阶段:预处理阶段按固定大小分段读取上下文,用测试时学习的快权重记忆网络做层局部梯度更新;推理阶段只从 query token 生成 key/value,不再重读任何上下文 token。推理代价与上下文长度解耦:预处理 O(N)、查询 O(1),峰值 GPU 显存不随上下文长度增长。128K token 下相比直接上下文内学习计算量与峰值显存各降约 80%,参数开销仅 6.4%,并在 RULER 针堆取数与词提取基准上保持强势。

为什么值得关注

冻结骨干上的模块化神经记忆:O(1) 查询、峰值显存不随上下文增长,128K 下省 80% 算力

English Abstract

We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

Obsidian 证据

  • 来源 digest: 论文流水线 2026-08-20(评分 8.5)。
  • 元数据与摘要经 opencli arxiv paper 核对;中文导读锚定摘要陈述的事实与数字。