Agent 与自动化 5.0 · 必读 2026-08-27 · 论文

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

自主循环 agent 的安全护栏普遍定义在单条轨迹上轨迹结束即重置论文证明这是组合性失败而非实现细节:面对把证据拆散到多次迭代里的攻击,任何轨迹窗口内监控器的真阳率都等于其假阳率(分离定理),而保留跨迭代状态的监控器可以完全分开干净与攻击轨迹;几何衰减的风险分数也不够,耐心对手的冷却期是常数不随水平线 N 增长提出的 LoopHarness 在循环层维护持久非衰减安全状态,在 mediated commits 与仲裁检测下限 _M 下,把未授权不可逆动作的期望次数压到与 N 无关的常数 B+m-1+m/_M,其中 B+m-1 项由 model-free 规则决定即使验证器完全合谋也成立附 Agent-SafetyBench 配对评测协议与跨迭代证据攻击套件

打开原文回到归档

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

External-scan entry · 20260830 · awesome-ai-field-notes
  • URL: https://arxiv.org/abs/2608.27141
  • PDF: https://arxiv.org/pdf/2608.27141v1
  • Authors: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
  • Published: 2026-08-27
  • Categories: cs.CR, cs.AI
  • Comment: -
  • Category: agents
  • Tags: agent-safety, autonomous-loops, monitoring, separation-result, loopharness
  • Quality Score: 5

中文摘要

自主循环 agent 的安全护栏普遍定义在单条轨迹上、轨迹结束即重置。论文证明这是组合性失败而非实现细节:面对把证据拆散到多次迭代里的攻击,任何轨迹窗口内监控器的真阳率都等于其假阳率(分离定理),而保留跨迭代状态的监控器可以完全分开干净与攻击轨迹;几何衰减的风险分数也不够,耐心对手的冷却期是常数、不随水平线 N 增长。提出的 LoopHarness 在循环层维护持久非衰减安全状态,在 mediated commits 与仲裁检测下限 δ_M 下,把未授权不可逆动作的期望次数压到与 N 无关的常数 B+m-1+m/δ_M,其中 B+m-1 项由 model-free 规则决定、即使验证器完全合谋也成立。附 Agent-SafetyBench 配对评测协议与跨迭代证据攻击套件。

English Abstract

Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。