Agent 与自动化 4.0 · 优秀 2026-08-27 · 论文

The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions

Cautious Bench 是第一个把 over-safety(误拒)作为 agent 护栏核心测量对象的 benchmark:756 对可判定的良性/孪生样本,每对在三种对象名下各测一遍(2268 对实测),另附 40 对不可判定样本单独报告构建期有一个门禁对每个样本按声明的授权策略机械重推标签,使标签成为策略的机械推论而非标注员的逐条判断,构成理想护栏的决策边界参照对五种设计的六个护栏实测后,全部出现 name-superstition 效应:同一个已授权动作,对象名字看起来吓人时被拒得更多对照实验里只有名字在变,偏差只能归因于护栏读了表面标签而没有读授权上下文做 agent 权限系统需要这类参照:要测误拒率,先要有理想护栏的决策边界图

打开原文回到归档

The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions

摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
  • Cautious Bench 是第一个把 over-safety(误拒)作为 agent 护栏核心测量对象的 benchmark:756 对可判定的良性/孪生样本,每对在三种对象名下各测一遍(2268 对实测),另附 40 对不可判定样本单独报告。构建期有一个门禁对每个样本按声明的授权策略机械重推标签,使标签成为策略的机械推论而非标注员的逐条判断,构成理想护栏的决策边界参照。对五种设计的六个护栏实测后,全部出现 name-superstition 效应:同一个已授权动作,对象名字看起来吓人时被拒得更多——对照实验里只有名字在变,偏差只能归因于护栏读了表面标签而没有读授权上下文。做 agent 权限系统需要这类参照:要测误拒率,先要有理想护栏的决策边界图。

论文信息

Abstract(原文)

Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one that maps the decision boundary of an ideal guardrail. Harvesting such a benchmark from real data is impractical: boundary cases are hard to collect, their labels hard to verify. The gap is real, so we construct Cautious Bench, the first benchmark to make over-safety the construct for agent guardrails; it codesigns each sample and its label with a stated authorization policy. A build-time gate re-derives every example to certify it, so each label is a mechanical consequence of the policy rather than an annotator's per-sample verdict, a reference against which researchers can measure real guardrails. The benchmark renders 756 Decidable benign/twin pairs, each under three object-name types (2,268 measured pairs), and 40 Undecidable pairs reported separately. Measuring six guardrails from five designs, we find a name-superstition effect: each over-refuses an authorized action more often under a scary-looking object name than a benign one. Since only the object name varies in the aforementioned contrast experiments, the deviation is the name's doing: the guardrails read the surface label, not the authorization context.

核心要点(英文摘要的中文提炼)

  • 论文题为 The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions,发表于 arXiv(cs.CR,2026-08-27 提交/更新)。
  • 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。

Obsidian 证据摘录

入选自 Obsidian《论文流水线 · 2026-08-31》速报第5篇:护栏误拒 benchmark,name-superstition 效应。