AI 编程 4.0 · 优秀 2026-08-26 · 文章

Breaking Claude Code Opus 5 Auto Mode

Johann Rehberger 演示一个简单的总结这个网站请求即可在 Auto Mode 下劫持 Claude Code Opus 5 并实现代码执行,小样本下攻击成功率 60-80%;而 Anthropic 委托第三方(Trajectory Labs)对 72 个间接注入场景各测十次的评估显示 Opus 5 Auto Mode 的注入得手率为 0.00%Auto Mode 8 月中旬起成为 Claude Code 默认模式,用安全分类器替代人工审批作者的核心提醒:Auto Mode 不能替代隔离环境运行与行为监控攻击链路利用下载解压 zip 再 import base64 时顺带导入同目录恶意 struct.py 的方式绕过防线...

打开原文回到归档

Breaking Claude Code Opus 5 Auto Mode

摘要中文导览(来自条目评分时的双语摘要,基于原文提炼):
  • Johann Rehberger 演示一个简单的“总结这个网站”请求即可在 Auto Mode 下劫持 Claude Code Opus 5 并实现代码执行,小样本下攻击成功率 60-80%;而 Anthropic 委托第三方(Trajectory Labs)对 72 个间接注入场景各测十次的评估显示 Opus 5 Auto Mode 的注入得手率为 0.00%。Auto Mode 8 月中旬起成为 Claude Code 默认模式,用安全分类器替代人工审批。作者的核心提醒:Auto Mode 不能替代隔离环境运行与行为监控。攻击链路利用下载解压 zip 再 import base64 时顺带导入同目录恶意 struct.py 的方式绕过防线;更值得警惕的是 Claude 自己检测到 compromise 并尝试清理恶意进程时,自动模式反而拦下了清理指令——安全机制本身成了失败的一部分。

文章信息

原文摘录(开头)

In this post, we explore how a simple website summary request hijacks Claude Code Opus 5 in Auto Mode and achieves code execution with 60-80% attack success rate using a small sample size.

This is interesting because a third-party evaluation commissioned by Anthropic showed a 0.00% prompt injection attack success rate for Opus 5 in Auto Mode.

Auto Mode replaces human approval prompts with a safety classifier. Since mid-August it is the default starting mode for Claude Code.

To make my key point right away: If you care about what’s happening and are worried about misalignment, hallucinations and prompt injection, then Auto Mode IS NOT a substitute for running your agent in an isolated environment and monitoring what it is up to.

Boris Cherny from Anthropic recently posted that layered defenses could reduce indirect prompt injection on unseen attacks to approximately zero. The layers were model training, input probes and an intent classifier. They hired a vendor (Trajectory Labs) to test 72 indirect prompt injection scenarios ten times each. The evaluation seems to not have a published benchmark name, and the shared chart shows **0.00% attack success for Opus 5 in A

Summary (EN)

Johann Rehberger demonstrates that a simple summarize-this-website request hijacks Claude Code Opus 5 in Auto Mode with a 60-80% attack success rate on a small sample — against Anthropic's commissioned third-party evaluation reporting 0.00% injection success for Opus 5 Auto Mode across 72 indirect prompt injection scenarios. Auto Mode (a safety classifier replacing human approval) became the default in mid-August. The exploit chain uses a download-unzip-then-import-base64 pattern that side-loads a malicious struct.py. Most striking: when Claude itself detected the compromise and tried to clean up the malicious process, Auto Mode blocked the cleanup — the safety mechanism became part of the failure.

Obsidian 证据摘录

入选自 Obsidian《AK RSS Digest · 2026-08-31》第4篇:Simon Willison 转载的 Johann Rehberger 攻击研究。