When Context Gets Root: Privilege Escalation in LLM Harnesses
External-scan entry · 20260830 · awesome-ai-field-notes
- URL: https://arxiv.org/abs/2608.27299
- PDF: https://arxiv.org/pdf/2608.27299v1
- Authors: Xingbang He, Yuanwei Chen, Yi Qian, Haiyang Wei, Ligeng Chen, Zenan Fu, Linzhang Wang, Hao Wu, Bing Mao
- Published: 2026-08-27
- Categories: cs.CR, cs.SE
- Comment: -
- Category: agents
- Tags: agent-security, instruction-hierarchy, privilege-escalation, coding-agents, harness
- Quality Score: 5
中文摘要
提出 instruction privilege escalation:agent harness 每次调用模型时构造上下文,可能把低权限内容抬升到更高 instruction level,让模型执行原本会拒绝的指令。作者用多 agent 机制在 6 个编码 agent harness 上实现 13 个攻击目标,覆盖机密性、完整性、可用性与远程代码执行: unrestricted 执行下全部 13 目标在全部 6 个 harness 得手,自动权限审查模式下 3 个提供该模式的 harness 也全部失守;还用 harness 自带的 persistent goals 与 scheduled tasks 复现了漏洞。对做 coding agent 权限模型的团队,这是一篇必须对照自查的威胁模型论文。
English Abstract
Instruction hierarchy is a model-side defense that assigns instructions different levels of privilege according to their sources. These levels constrain which content may direct model behavior. During agent execution, however, agent harnesses construct context for each model invocation. This construction can elevate low-level content to a higher instruction level and grant it greater model-facing privilege. We introduce instruction privilege escalation. In this attack, an attacker induces an agent to elevate low-level malicious content to a higher instruction level. The elevated content then causes the agent to execute instructions it would not follow at their original level. We evaluate this threat by using multi-agent mechanisms to achieve 13 attack objectives across six coding-agent harnesses. These objectives span confidentiality, integrity, availability, and remote code execution. With unrestricted action execution, the attacks achieve all 13 objectives on all six harnesses. Under automatic permission review, the attacks achieve all 13 objectives on all three harnesses that provide this mode. We further reproduce the vulnerability using harness-provided persistent goals and scheduled tasks. These results demonstrate the generality of instruction privilege escalation.
注:本文件为 external-scan cron 写入的 source body;如需更深入精读,请由 content-fetcher 任务补充完整正文。