Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
- ID: 13a8ad46
- 原文链接: https://arxiv.org/abs/2605.23989
- PDF: https://arxiv.org/pdf/2605.23989v1
- 作者: Jinhu Qi, Muzhi Li, Jiahong Liu, Yuqin Shu, Dianzhi Yu, Shicheng Ma, Wenqian Cui, Yiyang Zhao, Yiyi Chen, Ruoxi Jiang, Irwin King, Zenglin Xu
- 日期: 2026-05-17
- 更新: 2026-05-17
- 分类: agents
- 来源类型: paper
- 标签: survey, trustworthy-ai, safety, security, privacy, robustness, agentic-ai
- 质量评分: 5/5
- 抓取时间: 2026-07-30T04:18:47Z
中文导读
全面综述 agentic AI 的信任性,聚焦两个高风险部署维度:安全与鲁棒性隐私与系统安全每个维度按 agent 工作流梳理风险点,并给出阶段性缓解策略统一评测框架同时纳入 outcome 和 process 信号(约束违规/轨迹完整性/对抗成功率),提供场景到指标的发版决策指引开放挑战包括自进化 agent运行时监控/验证隐私保护个性化信任-效用权衡含开源 agentic 系统真实安全失败案例
为什么值得关注
高风险部署必读:从安全/鲁棒性/隐私/系统安全四维梳理 agentic AI 信任性,附统一评测框架 + 真实失败案例
- arXiv comment: 36 pages, 4 figures. Survey/review article on trustworthy agentic AI. Published in Academia AI and Applications, 2026
关键信息
- 论文标题:Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
- 作者:Jinhu Qi, Muzhi Li, Jiahong Liu, Yuqin Shu, Dianzhi Yu, Shicheng Ma, Wenqian Cui, Yiyang Zhao, Yiyi Chen, Ruoxi Jiang, Irwin King, Zenglin Xu
- arXiv:https://arxiv.org/abs/2605.23989
- 发布时间:2026-05-17
- arXiv 分类:cs.AI, cs.CL, cs.CR
- 关联标签:survey, trustworthy-ai, safety, security, privacy, robustness, agentic-ai
English Abstract
Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
English Summary
A comprehensive survey of trustworthy agentic AI examining two core dimensions critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, the authors clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. They consolidate evaluation into a unified metrics-and-benchmarks hub emphasizing both outcome and process signals (constraint violations, trace completeness, adversarial success rates) with scenario-to-metric guidance for release gating. Open challenges include self-evolving agents, runtime monitoring/verification, privacy-preserving personalization, and the trust-utility trade-off. Includes a case study of real-world security failures in open-source agentic systems.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。