SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control
External-scan entry · 20260830 · awesome-ai-field-notes
- URL: https://arxiv.org/abs/2608.27234
- PDF: https://arxiv.org/pdf/2608.27234v1
- Authors: Dylan Girrens, Guangjing Wang
- Published: 2026-08-27
- Categories: cs.CR, cs.SE
- Comment: Submitted to USENIX Security 2027. 24 pages, 17 figures
- Category: agents
- Tags: agent-security, information-flow-control, prompt-injection, persistence, usenix
- Quality Score: 5
中文摘要
针对持久化 LLM agent 跨查询的攻击面(攻击者数据改变控制流、进入敏感工具参数、污染后续查询),提出 plan-first 架构:planner 每查询只调用一次、以声明式 DSL 生成完整可执行计划,再用 dual-lattice 信息流控制同时跟踪机密性与完整性,覆盖显式数据流与控制依赖;持久化以带标签 artifact 存储,后续规划只暴露语义元数据、不再把不可信 payload 送回 planner。在 AgentDojo 上 tool_knowledge 攻击成功率降到 0,多查询扩展 AgentDojo-MQ 上降到 0.2%,作者同时指出严格完整性 enforcement 带来 security-utility tradeoff。投稿 USENIX Security 2027。
English Abstract
Large language model (LLM) agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources. Existing defenses typically protect either planning or individual tool interactions, but persistent agents face a broader threat: attacker-controlled data can alter control flow, enter security-sensitive tool arguments, or compromise later queries. We present SPA, a plan-first architecture that secures planning, execution, and cross-query state reuse. SPA invokes the planner once per query to generate a complete executable plan in a declarative domain-specific language, then applies dual-lattice information-flow control to track confidentiality and integrity across explicit data flows and control dependencies. To support persistence without re-exposing untrusted payloads to the planner, SPA stores execution results as labeled artifacts and reveals only semantic metadata during later planning. We evaluate SPA on AgentDojo and AgentDojo-MQ, which is our multi-query extension for measuring secure state reuse and delayed attacks. Under the 'tool_knowledge' attack, SPA with information-flow control reduces attack success to zero on AgentDojo and 0.2% on AgentDojo-MQ. Our results show that plan-first execution combined with label-preserving persistence can substantially strengthen persistent LLM agents, while revealing an important security-utility tradeoff introduced by strict integrity enforcement.
注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。