MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
- ID: 758b87c8
- 原文链接: https://arxiv.org/abs/2608.24189
- PDF: https://arxiv.org/pdf/2608.24189v1
- 作者: Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara
- 日期: 2026-08-25
- 更新: 2026-08-25
- 分类: agents
- 来源类型: paper
- 标签: memory, evaluation, conversational-ai, benchmark, arxiv
- 质量评分: 4/5
中文导读
记忆系统长期按 Direct QA 打分——能否从过去对话里召回事实 X。作者把这套打分放进 4 个月、40 名用户、1,872 次会话、7 种记忆条件的真实部署里检验:各条件 Direct QA 准确率从 19.7% 到 70.1% 拉开巨大,用户满意度却不变。解释是两者量的能力不同:基准测「问了才想起来」,对话需要「察觉相关就自然织进回复」。据此提出 MemUse:用部署中用户真实触发的记忆时刻、按整合质量打分——同一系统 Direct QA 拿 78.8%,在对话里实际引用这些事实的比例只有 7.9%,差 71 个百分点;且自然整合与满意度相关、Direct QA 不相关。EMNLP 2026 主会论文,部署语料与判分提示词开源。
为什么值得关注
MemUse:Direct QA 78.8% 的记忆系统对话中只自然引用 7.9%——检索能力与整合能力脱钩 71 个百分点,记忆验收指标要换
English Abstract
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.
Obsidian Notes
- Metadata and abstract fetched via
opencli arxiv paper 2608.24189 -f json(2026-08-27); response parsed list-or-dict tolerant. - EMNLP 2026 主会;开源仓库 github.com/ryuichi-sumida/memuse(摘要内给出)。
- 中文导读与价值判断锚定在论文摘要上,未补充摘要之外的实验细节。