Agent 与自动化 4.0 · 优秀 2026-08-10 · 论文

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

SHE 提出针对 LLM Agent 安全的"安全 Harness 演化"框架:将 harness 拆解为 System PromptRule BankSafety MemoryTool Policy 四个独立职责单元,基于 rollout 轨迹做归因分析并在结构化诊断驱动下做局部演化,在 Agent-SafetyBench 上相对静态 SafeHarness 实现 3.1 ASR 下降,并在 held-out AgentHarm 上泛化可零成本迁移到其他 Agent 模型

打开原文回到归档

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

  • ID: 8204e4c5
  • Original: https://arxiv.org/abs/2608.09885
  • PDF: https://arxiv.org/pdf/2608.09885v1
  • Authors: Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv Categories: cs.AI, cs.CV
  • AAIF Category: agents
  • Source Type: paper
  • Tags: agent-safety, llm-agent, harness, evolution, attribution
  • Quality Score: 4/5
  • Fetched At: 2026-08-12T04:29:18.480369+00:00

Chinese Guide

SHE 提出针对 LLM Agent 安全的"安全 Harness 演化"框架:将 harness 拆解为 System PromptRule BankSafety MemoryTool Policy 四个独立职责单元,基于 rollout 轨迹做归因分析并在结构化诊断驱动下做局部演化,在 Agent-SafetyBench 上相对静态 SafeHarness 实现 3.1 ASR 下降,并在 held-out AgentHarm 上泛化可零成本迁移到其他 Agent 模型

Why It Matters

SHE 提出针对 LLM Agent 安全的"安全 Harness 演化"框架:将 harness 拆解为 System PromptRule BankSafety MemoryTool Policy 四个独立职责单元...

This content page is grounded in the existing AAIF entry plus arXiv metadata fetched with opencli arxiv paper. It does not add experimental claims beyond the arXiv abstract and entry summary.

Key Metadata

  • Paper title: SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
  • Authors: Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
  • arXiv: https://arxiv.org/abs/2608.09885
  • PDF: https://arxiv.org/pdf/2608.09885v1
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv categories: cs.AI, cs.CV
  • Comment: Project: https://github.com/RainbowQTT/SHE
  • Related AAIF tags: agent-safety, llm-agent, harness, evolution, attribution

English Abstract

The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE decomposes the harness into four artifacts with explicit safety responsibilities, including the System Prompt, Rule Bank, Safety Memory, and Tool Policy, defining clear functional boundaries for localized evolution. Based on this decomposition, SHE introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation. Experiments on Agent-SafetyBench demonstrate that SHE effectively enhances safety through harness evolution, achieving a 3.1x ASR reduction compared with static SafeHarness, while also improving benign utility. The evolved harness further generalizes to unseen risks on the held-out AgentHarm benchmark and transfers across agent models without additional evolution.

English Summary

Safety Harness Evolution (SHE) treats the LLM agent harness the layer managing context, memory, tools, permissions, and runtime control as an evolvable artifact rather than fixed deployment infrastructure. It decomposes the harness into four responsibility-isolated pieces (System Prompt, Rule Bank, Safety Memory, Tool Policy), then runs an attribution-guided loop that turns trajectory failures into artifact-specific boundary refinements with safety-utility validation. On Agent-SafetyBench the evolved harness achieves a 3.1 ASR reduction versus the static SafeHarness baseline while improving benign utility; the harness also generalizes to the held-out AgentHarm benchmark and transfers across agent backbones without re-evolution.

Obsidian Notes

  • Generated by the AAIF content-fetcher cron mode from arXiv metadata and the existing AAIF entry.
  • Chinese guide and relevance notes are anchored to the entry summary, arXiv abstract, authors, dates, categories, and links.
  • Canonical content path is content/{entry_id}.md under the repository root; openclaw/content/ is intentionally avoided.