RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
- ID: 9f65064d
- 原文链接: https://arxiv.org/abs/2605.18805
- PDF: https://arxiv.org/pdf/2605.18805v1
- 作者: Imad Aouali, Flavian Vasile, Otmane Sakhi, Alexandre Gilotte, Benjamin Heymann
- 日期: 2026-05-11
- 更新: 2026-05-11
- 分类: agents
- 来源类型: paper
- 标签: recommendation-agents, agent-evaluation, benchmark, shopping-agents, cs.ir
- 质量评分: 4/5
- 抓取时间: 2026-07-25T04:19:47+00:00
中文导读
RecoAtlas 针对 LLM 推荐 Agent 生成推荐集合+自然语言解释的新形态,指出现有评估过度依赖小候选集重排或语义合理性摘要提出结合历史交互指标相关性/互补性/多样性效用代理,以及解释质量和语义一致性分开度量的 benchmark 与工具包
为什么值得关注
推荐 Agent 的评估要从解释听起来合理转向集合级效用互补性和多样性
这篇论文属于 AAIF 的 Agent 安全/评估线索:它不是泛泛讨论“让模型更安全”,而是把风险落到工具调用、权限继承、授权边界或集合级效用等可被工程系统观测与约束的对象上。对于正在构建工具型 Agent、连接器、沙箱评测或推荐型 Agent 的团队,它提供了一个更具体的检查维度:模型输出之外,系统还需要回答“Agent 被允许做什么、实际做了什么、这些动作对用户账户/环境/推荐集合造成什么后果”。
关键信息
- 论文标题:RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
- 作者:Imad Aouali, Flavian Vasile, Otmane Sakhi, Alexandre Gilotte, Benjamin Heymann
- arXiv:https://arxiv.org/abs/2605.18805
- 发布时间:2026-05-11
- arXiv 分类:cs.IR, cs.AI, cs.LG
- 关联标签:recommendation-agents, agent-evaluation, benchmark, shopping-agents, cs.ir
English Abstract
LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet existing evaluations often reduce this setting to reranking small shortlisted candidate sets or judge reports mainly by semantic plausibility. We introduce Recommendation Atlas (Agentic Tool-Level Assessment for Shopping), or RecoAtlas, a benchmark and toolkit for evaluating shopping agents with behavior-grounded metrics. RecoAtlas complements held-out interaction metrics with learned utility proxies for relevance, complementarity, and diversity derived from interaction data, while separately measuring semantic coherence and explanation quality. Its controlled tool environment exposes agents to either semantic, behavior-aligned, or faulty tools, enabling diagnosis of whether performance gains arise from stronger reasoning, better signals, or more effective tool-use policies. Across controlled experiments, we show that RecoAtlas exhibits key properties of a meaningful benchmark for agentic systems: performance scales with model capacity and test-time compute, improves with stronger and better-aligned tools, degrades under noisy or misaligned signals, and reveals that semantic plausibility does not necessarily capture behavior-grounded utility. RecoAtlas provides a foundation for developing and evaluating shopping assistants that optimize not only for plausible recommendations, but also for coherent, behaviorally grounded recommendation sets.
English Summary
LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet existing evaluations often reduce this setting to reranking small shortlisted candidate sets or judge reports mainly by semantic plausibility. We introduce Recommendation Atlas (Agentic Tool-Level Assessment for Shopping), or RecoAtlas, a benchmark and toolkit for evaluating shopping agents with behavior-grounded metrics. RecoAtlas complements held-out interaction metrics with learned utility proxies for relevance, complementarity, and diversity derived from interaction data, while separately measuring semantic coherence and explanation quality. Its controlled tool environment exposes agents to either semantic, behavior-aligned, or faulty tools, enabling diagnosis of whether performance gains arise from stronger reasoning, better signals, or more effective tool-use policies. Across controlled experiments, we show that RecoAtlas exhibits key properties of a meaningful benchmark for agentic systems: performance scales with model capacity and test-time compute, improves with stronger and better-aligned tools, degrades under noisy or misaligned signals, and reveals that semantic plausibility does not necessarily capture behavior-grounded utility. RecoAtlas provides a foundation for developing and evaluating shopping assistants that optimize not only for plausible recommendations, but also for coherent, behaviorally grounded recommendation sets.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。