What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
- ID: 8aad2976
- 原文链接: https://arxiv.org/abs/2608.16852
- PDF: https://arxiv.org/pdf/2608.16852
- 作者: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
- 日期: 2026-08-17
- 更新: 2026-08-17
- 分类: learning
- 来源类型: paper
- 标签: compliance, safety, audit, activation-probes, guard-models, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-19T04:47:43Z
中文导读
对部署侧 LLM 合规监测器的审计发现'规则盲视'(rule blindness):删除置换或替换适用规则,所有被测 guard 与激活探针的检测准确率几乎不变包括能正确引用适用条款的策略条件 guard,把条款换成宽松版后判定也几乎不动作者构造规则x场景交叉基准(单看任一维度都无法预测标签),在此前基准未排除的设计下证实该失效;被测对象中只有逐步推理而非任何快速检测器能逃出为规模化审计另提出免训练的激活层检测器 Internal Compliance Score(ICS)对把合规检测当作法律/审计控制直接使用的做法是一记直接警示
为什么值得关注
合规检测器'规则盲视'审计:换掉规则判定不变,只有逐步推理能逃出该失效
关键信息
- 论文标题: What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
- 作者: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
- arXiv: https://arxiv.org/abs/2608.16852
- 发布时间: 2026-08-17
- arXiv 分类: cs.AI
- 关联标签: compliance, safety, audit, activation-probes, guard-models, arxiv
English Abstract
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.
English Summary
Regulatory compliance monitoring of deployed language models is increasingly implemented as a legal and audit control, but it is meaningful only if the detector's verdict depends on the stated rule rather than surface features. The audit shows this condition fails across the current class of compliance detectors - a failure the authors call rule blindness: deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe tested, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when the clause is swapped for its permissive counterpart....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- metadata source: opencli