DolphinBench: Mapping the Pareto Frontier of Agent Memory
- ID: fcc76d41
- 原文链接: https://arxiv.org/abs/2609.24971
- PDF: https://arxiv.org/pdf/2609.24971v2
- 作者: Soumil Rathi, Deshraj Yadav, Taranjeet Singh
- 日期: 2026-09-21
- 更新: 2026-09-22
- 分类: agents
- 来源类型: paper
- 标签: arxiv, agent-memory, benchmark, evaluation, paper
- 质量评分: 4/5
- 抓取时间: 2026-09-23T04:27:36+00:00
中文导读
现有 agent 记忆评测多为对话问答格式:问题本身已经提示需要检索哪条事实,而且大多只看准确率,记忆系统可以用不合理的成本/延迟换分DolphinBench 改用任务完成度直接评记忆:3 个知识工作 persona每个约 50 万 token 的用户消息历史,每 persona 200 个依赖历史信息的任务;每个任务都用带相关历史能成功 + 不带历史会失败的双跑验证评测强制同时上报总成本延迟与准确率,让记忆系统可以在准确率-成本-延迟的整体权衡下比较;作者称现有记忆 benchmark 没有同时做到这三点的数据集与评测代码在 dolphinbench.ai
为什么值得关注
DolphinBench:50 万 token 级 persona 历史 + 双跑验证 + 强制上报成本/延迟,评 agent 记忆不再只看准确率
关键信息
- 论文标题:DolphinBench: Mapping the Pareto Frontier of Agent Memory
- 作者:Soumil Rathi, Deshraj Yadav, Taranjeet Singh
- arXiv:https://arxiv.org/abs/2609.24971
- 发布时间:2026-09-21
- arXiv 分类:cs.CL, cs.AI
- 评论/页数:6 pages, 2 figures
- 关联标签:arxiv, agent-memory, benchmark, evaluation, paper
English Abstract
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.
English Summary
Most agent-memory benchmarks use a conversational QA format where the question itself signals that a fact must be retrieved and often which one, and they grade accuracy alone, letting memory systems buy scores with unreasonable cost/time tradeoffs. DolphinBench evaluates memory directly through an agent's task completion: three knowledge-work personas with roughly 500k tokens of user messages each, and 200 tasks per persona that depend on information from that history. Every task is verified by running an agent with and without the relevant history, requiring success with it and failure without it. All evaluations must report total cost and latency alongside accuracy, enabling holistic evaluation of memory systems; the authors state no existing memory benchmark combines all three. Dataset and evaluation code are at dolphinbench.ai.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。