The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- 同一份泄露指令在六个模型上以十类直白注入形式全部被拒(gpt-4o 得手率 0%),但把它改写成“完整性签名”“配置字段”“仿冒可信域名”等框架后,gpt-4o 的得手率从 0% 抬升到 100%。攻击成本分三级:对已知机制改 3 个措辞即有 96% 成功率;在已知模板中换一个字段最高 60%;围绕新机制从零写一整页只有 0/130——可复用的攻击资产是模板而非机制。消融实验确认机制是指令/数据混淆而非对齐被攻破:删除保密策略后基础攻击归零,框架化攻击仅从 31.9% 升至 38.1%。有效防御都在约束出口:目的地白名单在目的地封闭时得手率 0%,planner/reader 能力隔离 0%;SecAlign 微调在工具 agent 上仅剩 32.5% 防御效果,输出规范化护栏被一次 ROT13 编码完全绕过。
论文信息
- arXiv ID: 2608.27092
- 作者: Md Habibur Rahman, Jaeho Kim
- 发表: 2026-08-27(更新:2026-08-27)
- 分类: cs.CR
- 原文链接: https://arxiv.org/abs/2608.27092
- PDF: https://arxiv.org/pdf/2608.27092
- 标签:
prompt-injectionagent-securityexfiltrationdefensestool-using-agents
Abstract(原文)
A tool-using LLM agent that reads attacker-controlled web content while holding a secret faces indirect prompt injection: the content may make it exfiltrate the secret. In a safe synthetic lab (canary secret, mock tools, matched clean-vs-poisoned metric) we report the framing gap: across six models, ten overt injection classes are refused (gpt-4o 0%), but reframing the identical leak as a mandatory integrity signature, config field, or look-alike "trusted" host drives gpt-4o 0% to 100%. The attack is cheap, and its cost is three-level: paraphrasing a known mechanism is trivial (96% at 3 wordings), swapping the field inside a known-effective template is also cheap (up to 60%), while authoring a fresh page around a new mechanism is hard (0/130) -- the reusable asset is the template, not the mechanism. An ablation shows the mechanism is instruction/data confusion, not defeated alignment: removing the confidentiality policy leaves base attacks at 0% and moves reframing only 31.9% to 38.1%. What closes the gap is payload-blind checks: a destination allow-list (0%, when destinations are closed) and a capability-isolating planner/reader split (0%). A broad "in any form" policy clause also closes it at the acting model (to 0%) but is brittle (dropping the catch-all reopens it to 48.8%). A published fine-tuning defense (SecAlign, CCS 2025) does not close it on a tool agent (32.5%, positive-control-validated), nor does channel separation (38.8%); an output-normalizing guard loses to a held-out encoding (ROT13, 100%). Robustness comes from constraining the destination or isolating the capability, not from the acting model recognizing the attack.
核心要点(英文摘要的中文提炼)
- 论文题为 The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents,发表于 arXiv(cs.CR,2026-08-27 提交/更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
Obsidian 证据摘录
入选自 Obsidian《论文流水线 · 2026-08-31》速报第3篇:注入框架化 0% 到 100%,防御对照数字密。