MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
- ID: ea34da35
- 原文链接: https://arxiv.org/abs/2603.25973
- PDF: https://arxiv.org/pdf/2603.25973v1
- 作者: Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang, Zheng Hui, Chen Wang, Michelle Gong, Philip S. Yu
- 日期 / 版本: Submitted on 2026-03-26
- 分类: agents
- 来源类型: paper
- arXiv 分类: cs.CL
- arXiv 备注: Published as a workshop paper in Lifelong Agent @ ICLR 2026
- 标签: llm-agent, memory, long-context, personalization, benchmark, 2603-25973
- 质量评分: 4/5
- 抓取时间: 2026-07-27T12:26:00.305549+00:00
中文导读
论文提出 MemoryCD,用于 benchmarking LLM Agents 的 long-context user memory 与 lifelong cross-domain personalization摘要主题表明其关注 Agent 是否能在长期多领域交互中保存并正确调用用户记忆,是个人助理与持久化记忆系统的重要评测线索
为什么值得关注
MemoryCD 评测 LLM Agent 的长期用户记忆,关注跨领域个性化场景里的记忆保持与调用
这篇内容值得放进 AAIF,是因为它围绕 Agent / Coding Agent 系统中的一个具体评测或训练问题展开;本页基于 arXiv 元数据、摘要与条目已有摘要整理,未补充摘要之外的实验细节。
关键信息
- 论文标题:MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
- 作者:Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang, Zheng Hui, Chen Wang, Michelle Gong, Philip S. Yu
- arXiv:https://arxiv.org/abs/2603.25973
- 发布时间 / 修订:Submitted on 2026-03-26
- arXiv 分类:cs.CL
- arXiv 备注:Published as a workshop paper in Lifelong Agent @ ICLR 2026
- 关联标签:llm-agent, memory, long-context, personalization, benchmark, 2603-25973
English Abstract
Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{MemoryCD}, the first large-scale, user-centric, cross-domain memory benchmark derived from lifelong real-world behaviors in the Amazon Review dataset. Unlike existing memory datasets that rely on scripted personas to generate synthetic user data, \textsc{MemoryCD} tracks authentic user interactions across years and multiple domains. We construct a multi-faceted long-context memory evaluation pipeline of 14 state-of-the-art LLM base models with 6 memory method baselines on 4 distinct personalization tasks over 12 diverse domains to evaluate an agent's ability to simulate real user behaviors in both single and cross-domain settings. Our analysis reveals that existing memory methods are far from user satisfaction in various domains, offering the first testbed for cross-domain life-long personalization evaluation.
English Summary
Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{MemoryCD}, the first large-scale, user-centric, cross-domain memory benchmark derived from lifelong real-world behaviors in the Amazon Review dataset. Unlike existing memory datasets that rely on scripted personas to generate synthetic user data, \textsc{MemoryCD} tracks authentic user interactions across years and multiple domains. We construct a multi-faceted long-context memory evaluation pipeline of 14 state-of-the-art LLM base models with 6 memory method baselines on 4 distinct personalization tasks over 12 diverse domains to evaluate an agent's ability to simulate real user behaviors in both single and cross-domain settings. Our analysis reveals th
Obsidian Notes
- 内容获取路径:优先尝试
opencli arxiv paper 2603.25973 -f json,本页使用返回的 arXiv 元数据与摘要回填。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上。
- 现代站点生成器按
content/{entry.id}.md查找内容页;本文件写入 canonical content 目录,而不是openclaw/content/。