Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes
- ID: 5b06cb82
- 原文链接: https://arxiv.org/abs/2608.20685
- PDF: https://arxiv.org/pdf/2608.20685v1
- 作者: Neeraj Yadav
- 日期: 2026-08-21
- 更新: 2026-08-21
- 分类: agents
- 来源类型: paper
- 标签: agent-memory, rag, code-assistant, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-25T15:40:00+00:00
中文导读
RAG 没有时间概念:函数改名、接口迁移、依赖升级之后,旧值和新值在向量空间里相似度几乎一样,检索系统无法判断哪个是当前事实。这篇把「覆盖式记忆」放到真实软件历史上验证:从 707 个 GitHub issue 提取 130 个干净单值状态变迁,MemStrata 拿到 0.91 答案准确率而 RAG 只有 0.57-0.59;被强制作答时 RAG 有 36-38% 概率给出已被取代的旧值,MemStrata 压到约 0 且延迟保持在检索量级。给代码 agent 做记忆系统的人建议精读。
为什么值得关注
RAG 没有时间概念:函数改名、接口迁移、依赖升级之后,旧值和新值在向量空间里相似度几乎一样,检索系统无法判断哪个是当前事实。 实验与数字均来自论文摘要本身。
关键信息
- 论文标题:Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes
- 作者:Neeraj Yadav
- arXiv:https://arxiv.org/abs/2608.20685
- 发布时间:2026-08-21
- arXiv 分类:cs.SE, cs.AI, cs.CL, cs.LG
English Abstract
Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint moves, a dependency is bumped - RAG retrieves both the old and new value with near-identical similarity and cannot tell which is current, so it serves the superseded value. Paper 1 showed, on synthetic single-value benchmarks, that a deterministic (subject, relation, object) supersession memory eliminates this failure. Here we validate it end-to-end on real software history. From 707 real GitHub issues (SWE-bench Lite + Verified) we extract 130 clean atomic state transitions, a fix that changes one identifiable value from a pre-fix to a post-fix form, and render each marker-free (the stale and current statements differ only in the value). On this set, MemStrata reaches 0.91 answer accuracy versus RAG's 0.57-0.59; and, the structural result, when forced to answer RAG serves the superseded value 36-38% of the time (an LLM reranker does not help) while MemStrata drives this to ~0, at RAG retrieval latency (~2.1 s vs ~18 s for the reranker). We are explicit about scope: only ~18% of real fixes are clean atomic transitions; Paper 2 isolates the memory mechanism on that class, and extraction coverage of the remaining fixes is the orthogonal problem we defer to follow-on work. A real product bug surfaced and was fixed during the study (a case/punctuation-insensitive value comparison), with the moat property (deterministic-supersession accuracy on clean code mutations) preserved and verified.
English Summary
Validates overwrite-style (subject, relation, object) memory on real software histories: from 707 GitHub issues, 130 clean single-value state transitions are extracted; MemStrata reaches 0.91 answer accuracy vs RAG's 0.57-0.59, drives forced-answer stale-value rates from 36-38% to ~0, and stays at retrieval-scale latency (~2.1s).
Obsidian Notes
- 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
- 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。