Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?
External-scan entry · 20260830 · awesome-ai-field-notes
- URL: https://arxiv.org/abs/2608.27443
- PDF: https://arxiv.org/pdf/2608.27443v1
- Authors: Ting Yan
- Published: 2026-08-27
- Categories: cs.HC, cs.CR
- Comment: 15 pages, 5 figures
- Category: agents
- Tags: agent-permissions, human-ai-interaction, user-study, overreach, consent
- Quality Score: 4
中文摘要
113 名非专业软件背景被试的对照实验,比较逐动作人工审批(HITL)、逐动作模型自动审查(AUTO)与用户预先写好的 allow/ask/never 后果策略(POLICY)三种 agent 权限机制。18 动作模拟日含 7 个越权动作:POLICY 比 HITL 少拦 20.1 个百分点、比 AUTO 少拦 14.5 个百分点的越权动作;运行时提示从 18.0 降到 10.9,但算上规则设置时间总干预时间并没有可靠下降。140 条规则里被试 114 条选了 ask,等于把越权决定又推回运行时;148 个被执行的越权动作中 133 个经人工批准。结论是偏好与承诺之间存在落差:反复选 ask 保留了逐案选择权,却让常设策略无法提前定案。
English Abstract
AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored "allow", "ask", or "never" rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions. POLICY blocked less overreach than HITL (-20.1 percentage points, 95% CI [-32.1, -8.1]) and AUTO (-14.5 points, 95% CI [-25.8, -3.2]). POLICY lowered runtime prompts from 18.0 to 10.9, but total intervention time was not reliably lower when rule setup was included. Exploratory analysis showed that participants chose "ask" for 114 of 140 POLICY rules, returning most overreach actions to runtime. Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under "allow" rules. Across all 7 overreach actions, POLICY had the highest approval rate. Counterintuitively, user-authored rules did not by themselves provide stronger protection: many actions outside users' original requests went through after users approved them. These results reveal a gap between preference and commitment: repeatedly choosing "ask" preserves case-by-case choice but prevents a standing policy from settling decisions in advance.
注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。