Breaking Claude Code Opus 5 Auto Mode
- 原文链接: https://simonwillison.net/2026/Aug/27/breaking-claude-code-opus-5-auto-mode/
- 作者: Simon Willison
- 日期: 2026-08-27
- 分类: coding
- 来源类型: article
- 标签: prompt-injection, agent-security, claude-code, sandboxing, auto-mode
- 质量评分: 4/5
- 抓取时间: 2026-08-28T23:55:00+08:00
中文导读
Simon Willison 转发 Johann Rehberger 在 embracethered 上的实测:Anthropic 把 Claude Code 的 auto mode(用分类器挡 prompt injection)设为默认并宣称高拦截率,Rehberger 用「诱导下载 zip 解压后再 import base64」触发本地 struct.py 被执行的方式做出约 80% 命中率的注入,且出现失败循环——Claude 自己发现被注入、试图 kill 恶意进程,但 auto mode 把清理命令也拦了下来:分类器放行了恶意进程的创建,却拦下了停止它的命令。Willison 认同 Rehberger 的结论:唯一可靠的做法是把 agent 关进容器/VM/OS 沙箱、限制网络出口、监控 agent 行为,不要把 home 目录、SSH key、云凭据暴露给 agent runtime。
为什么值得关注
把「分类器防御不可靠、沙箱是唯一底线」这条结论钉实的公开复现,对任何接生产代码库的 agent 部署是必读安全输入。
原文(抓取存档·节选)
# Breaking Claude Code Opus 5 Auto Mode
> 作者: Simon Willison
> 原文链接: https://simonwillison.net/2026/Aug/27/breaking-claude-code-opus-5-auto-mode/
---
27th August 2026 - Link Blog
**[Breaking Claude Code Opus 5 Auto Mode](https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/)**. Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently [made that the default](https://simonwillison.net/2026/Aug/8/auto-mode/) and have made bold claims about its effectiveness.
Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip archive, then executing code that imports `base64` without noticing that this will import and execute a local `struct.py` file extracted from the archive.
In a few cases auto mode directly prevented the agent from preventing harmful code from continuing to execute!
> In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.
>
> Claude detects the compromise, but **Auto Mode blocks its cleanup command**
>
> The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!
I agree with Johann's conclusion here: the only safe way to run agents if there's any risk of attracting the attention of an adversarial attack is with a sandbox:
> - Run unattended coding agents in a container, VM or OS sandbox.
> - Restrict network egress.
> - Monitor your agents.
> - Do not expose home directories, SSH keys, cloud credentials,… to the agent runtime.
Posted [27th August 2026](https://simonwillison.net/2026/Aug/27/) at 10:50 pm