Agent 与自动化 4.0 · 优秀 2026-09-04 · 论文

Does Your Agent's Memory Survive a Model Upgrade? Memory Portability Study

2609.05339 (cs.AI/cs.CL/cs.IR,9 月 4 日) 控制研究:同一段对话历史分别用 LC-RAW 长上下文RAG 分块NOTES 自然语言摘要KG-fixed 知识图谱四种记忆形态承载,在 48 条合成历史 + 随机答案码 + 两个 <10B 开源模型下 对比升级底座后的表现18 页 7 表,核心问题是"模型一换,智能体记忆还能不能用"

打开原文回到归档

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Source: https://arxiv.org/abs/2609.05339
Authors: Ankit Goyal, Jaideep Ray
Published: 2026-09-04
Categories: cs.AI, cs.CL, cs.IR
PDF: https://arxiv.org/pdf/2609.05339v1

Abstract

Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into natural-language notes (NOTES), or normalized into a fixed-schema knowledge graph (KG-fixed). The study uses 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models with sub 10 billion parameters. Our measurements show that fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap. Conversely, compressed NOTES exhibit high model coupling, with accuracy shifting asymmetrically by $+9.91$ or $-13.28$ percentage points depending on the specific migration direction. In RAG systems, partial embedding migrations using a 50/50 mixed index capture only a 4.96-point accuracy improvement, forfeiting the majority of the 11.90-point gain achieved through full re-embedding. Diagnostic decomposition attributes 80% ($0.467 \pm 0.014$) of the NOTES accuracy deficit to information lost during initial construction, whereas retrieval failures drive 81% ($0.364 \pm 0.012$) of the RAG deficit. Finally, store-only repair of NOTES fails to reach a 90% performance recovery target in all 48 test cases, whereas retaining the raw source history enables successful recovery in 34 of 48 cases for one tested direction. These findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair.

Key Findings

  • Setup: Same history preserved verbatim across four memory representations — LC-RAW (long-context), RAG chunks, NOTES (model-compressed natural-language notes), KG-fixed (fixed-schema knowledge graph). 48 synthetic histories with randomized answer codes, exact scoring, two open-weight sub-10B models.
  • Fixed schema transfers: After a writer-model swap, KG-fixed accuracy changes by only +0.0004 ± 0.0020 — the only representation that is essentially model-agnostic.
  • Compressed notes are model-coupled: NOTES accuracy shifts asymmetrically by +9.91 or −13.28 points depending on migration direction; 80% (0.467 ± 0.014) of the NOTES deficit comes from information lost at initial construction, not from migration itself.
  • RAG needs full re-embedding: A 50/50 mixed-embedding index recovers only 4.96 of the 11.90-point gain from full re-embedding; retrieval failures drive 81% (0.364 ± 0.012) of the RAG deficit.
  • Repair needs the raw history: Store-only repair of NOTES never reaches the 90% recovery target (0/48 cases); keeping the raw source history enables recovery in 34/48 cases for one tested direction.
  • Recommendations: Direction-specific migration testing, strict embedding-space isolation, and retention of raw source histories for memory repair.

中文概要

模型升级很频繁,但 agent 的记忆迁移很少有人系统研究:记忆库原封不动,换了底层模型照样可能"失忆"——新模型对旧笔记的解读不同、新旧 embedding 混用破坏检索、没有原始证据时修复失败。本文把同一份历史分别存成四种表示(长上下文原文 LC-RAW、RAG 分块、模型压缩的 NOTES 自然语言笔记、固定 schema 知识图谱 KG-fixed),用 48 条合成历史、随机化答案编码、精确打分和两个 10B 以下开源模型做受控对比。结论:固定 schema 的 KG 最抗迁移,换写入口径后准确率仅变化 +0.0004±0.0020;压缩 NOTES 与模型强耦合,迁移方向不同准确率不对称地 +9.91 或 −13.28 个百分点,且其 80% 的差距来自初次压缩时的信息损失;RAG 半量混合 embedding 只拿到全量重嵌入 11.90 点收益中的 4.96 点,81% 的差距来自检索失败;只对 NOTES 做存储侧修复在 48 例中无一达到 90% 恢复目标,而保留原始历史后一个方向上 34/48 例修复成功。实践建议:按方向做迁移测试、严格隔离 embedding 空间、保留原始 source history。