Agent 与自动化 4.0 · 优秀 2026-03-26 · 论文

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personal...

论文提出 MemoryCD,用于 benchmarking LLM Agents 的 long-context user memory 与 lifelong cross-domain personalization摘要主题表明其关注 Agent 是否能在长期多领域交互中保存并正确调用用户记忆,是个人助理与持久化记忆系统的重要评测线索

打开原文回到归档

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

  • ID: ea34da35
  • 原文链接: https://arxiv.org/abs/2603.25973
  • PDF: https://arxiv.org/pdf/2603.25973v1
  • 作者: Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang, Zheng Hui, Chen Wang, Michelle Gong, Philip S. Yu
  • 日期 / 版本: Submitted on 2026-03-26
  • 分类: agents
  • 来源类型: paper
  • arXiv 分类: cs.CL
  • arXiv 备注: Published as a workshop paper in Lifelong Agent @ ICLR 2026
  • 标签: llm-agent, memory, long-context, personalization, benchmark, 2603-25973
  • 质量评分: 4/5
  • 抓取时间: 2026-07-27T12:26:00.305549+00:00

中文导读

论文提出 MemoryCD,用于 benchmarking LLM Agents 的 long-context user memory 与 lifelong cross-domain personalization摘要主题表明其关注 Agent 是否能在长期多领域交互中保存并正确调用用户记忆,是个人助理与持久化记忆系统的重要评测线索

为什么值得关注

MemoryCD 评测 LLM Agent 的长期用户记忆,关注跨领域个性化场景里的记忆保持与调用

这篇内容值得放进 AAIF,是因为它围绕 Agent / Coding Agent 系统中的一个具体评测或训练问题展开;本页基于 arXiv 元数据、摘要与条目已有摘要整理,未补充摘要之外的实验细节。

关键信息

  • 论文标题:MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
  • 作者:Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang, Zheng Hui, Chen Wang, Michelle Gong, Philip S. Yu
  • arXiv:https://arxiv.org/abs/2603.25973
  • 发布时间 / 修订:Submitted on 2026-03-26
  • arXiv 分类:cs.CL
  • arXiv 备注:Published as a workshop paper in Lifelong Agent @ ICLR 2026
  • 关联标签:llm-agent, memory, long-context, personalization, benchmark, 2603-25973

English Abstract

Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{MemoryCD}, the first large-scale, user-centric, cross-domain memory benchmark derived from lifelong real-world behaviors in the Amazon Review dataset. Unlike existing memory datasets that rely on scripted personas to generate synthetic user data, \textsc{MemoryCD} tracks authentic user interactions across years and multiple domains. We construct a multi-faceted long-context memory evaluation pipeline of 14 state-of-the-art LLM base models with 6 memory method baselines on 4 distinct personalization tasks over 12 diverse domains to evaluate an agent's ability to simulate real user behaviors in both single and cross-domain settings. Our analysis reveals that existing memory methods are far from user satisfaction in various domains, offering the first testbed for cross-domain life-long personalization evaluation.

English Summary

Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{MemoryCD}, the first large-scale, user-centric, cross-domain memory benchmark derived from lifelong real-world behaviors in the Amazon Review dataset. Unlike existing memory datasets that rely on scripted personas to generate synthetic user data, \textsc{MemoryCD} tracks authentic user interactions across years and multiple domains. We construct a multi-faceted long-context memory evaluation pipeline of 14 state-of-the-art LLM base models with 6 memory method baselines on 4 distinct personalization tasks over 12 diverse domains to evaluate an agent's ability to simulate real user behaviors in both single and cross-domain settings. Our analysis reveals th

Obsidian Notes

  • 内容获取路径:优先尝试 opencli arxiv paper 2603.25973 -f json,本页使用返回的 arXiv 元数据与摘要回填。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上。
  • 现代站点生成器按 content/{entry.id}.md 查找内容页;本文件写入 canonical content 目录,而不是 openclaw/content/