Agent 与自动化 4.0 · 优秀 2025-02-13 · 论文

AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration

AgentGuard:复用智能体编排器自身能力做工具编排安全评测工具使用让被攻陷的智能体可执行后果更严重的恶意工作流;该框架让 LLM 编排器凭其工具功能知识可扩展的真实工作流生成能力与工具执行特权充当自身安全评估器,四阶段运作:发现不安全工作流真实执行中验证生成安全约束部署时约束智能体行为以达成基线安全保证

打开原文回到归档

AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration

  • ID: 457ee188
  • 原文链接: https://arxiv.org/abs/2502.09809
  • PDF: https://arxiv.org/pdf/2502.09809
  • 作者: Jizhou Chen, Samuel Lee Cong
  • 日期: 2025-02-13
  • 更新: 2025-02-13
  • 分类: agents
  • 来源类型: paper
  • 标签: tool-use, safety-evaluation, red-teaming, orchestrator, constraints, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-06 12:21 UTC

中文导读

AgentGuard:复用智能体编排器自身能力做工具编排安全评测工具使用让被攻陷的智能体可执行后果更严重的恶意工作流;该框架让 LLM 编排器凭其工具功能知识可扩展的真实工作流生成能力与工具执行特权充当自身安全评估器,四阶段运作:发现不安全工作流真实执行中验证生成安全约束部署时约束智能体行为以达成基线安全保证

为什么值得关注

用编排器评编排器:自主发现并真实执行验证不安全工具工作流,再生成部署期安全约束兜底

关键信息

  • 论文标题:AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
  • 作者:Jizhou Chen, Samuel Lee Cong
  • arXiv:https://arxiv.org/abs/2502.09809
  • 发布时间:2025-02-13
  • arXiv 分类:cs.CR, cs.AI
  • 关联标签:tool-use, safety-evaluation, red-teaming, orchestrator, constraints, arxiv

English Abstract

The integration of tool use into large language models (LLMs) enables agentic systems with real-world impact. In the meantime, unlike standalone LLMs, compromised agents can execute malicious workflows with more consequential impact, signified by their tool-use capability. We propose AgentGuard, a framework to autonomously discover and validate unsafe tool-use workflows, followed by generating safety constraints to confine the behaviors of agents, achieving the baseline of safety guarantee at deployment. AgentGuard leverages the LLM orchestrator's innate capabilities - knowledge of tool functionalities, scalable and realistic workflow generation, and tool execution privileges - to act as its own safety evaluator. The framework operates through four phases: identifying unsafe workflows, validating them in real-world execution, generating safety constraints, and validating constraint efficacy. The output, an evaluation report with unsafe workflows, test cases, and validated constraints, enables multiple security applications. We empirically demonstrate AgentGuard's feasibility with experiments. With this exploratory work, we hope to inspire the establishment of standardized testing and hardening procedures for LLM agents to enhance their trustworthiness in real-world applications.

English Summary

The integration of tool use into large language models (LLMs) enables agentic systems with real-world impact. In the meantime, unlike standalone LLMs, compromised agents can execute malicious workflows with more consequential impact, signified by their tool-use capability. We propose AgentGuard, a framework to autonomously discover and validate unsafe tool-use workflows, followed by generating safety constraints to confine the behaviors of agents, achieving the baseline of safety guarantee at deployment. AgentGuard leverages the LLM orchestrator's innate capabilities - knowledge of tool functionalities, scalable and realistic workflow generation, and tool execution privileges - to act as its own safety evaluator. The framework operates through four phases: identifying unsafe workflows, validating them in real-world execution, generating safety constraints, and validating constraint efficacy. The output, an evaluation report with unsafe workflows, test cases, and validated constraints, enables multiple security applications. We empirically demonstrate AgentGuard's feasibility with experiments. With this exploratory work, we hope to inspire the establishment of standardized testing and hardening procedures for LLM agents to enhance their trustworthiness in real-world applications.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。