AI 编程 4.0 · 优秀 2026-10 · 文章

Getting the most out of Opus 5.5 in Claude and Claude Code

Anthropic 官方指南Getting the most out of Opus 5.5 in Claude and Claude Code(Addy Osmani 执笔)出发点:Opus 5.5 能更长地独立工作更清楚地汇报自己做了什么并自行决定每次回答花多少思考对应建议:把整个任务连同完成的定义一次性交给模型,写明唯一需要停下来问人的情形;中途想起的事作为新消息补发而不是重启任务;从 promptCLAUDE.mdskills 里删掉think hard/carefully类指令(官方测试中删掉后更早开始作答且质量无明显下降);设计任务列出不要的风格清单比要现代简洁更有效...

打开原文回到归档

Getting the most out of Opus 5.5 in Claude and Claude Code

Content-fetcher entry · 2026-10-04 · awesome-ai-field-notes

中文摘要

Anthropic 官方 Opus 5.5 使用指南(Addy Osmani 执笔,2026-09-22 发布,9 分钟)。出发点:Opus 5.5 与前代行为不同——能更长时间独立工作、更直白地汇报自己做了什么、每次回复前自行决定思考量。因此整篇指南围绕三件事:怎么提需求、怎么驾驭长任务、怎么核查结果。核心建议:把整个任务连同"完成"的定义一次性交给模型,写明唯一需要停下来问人的情形;从 prompt、CLAUDE.md、skills 里删掉 think hard / think carefully 类指令——官方在聊天产品里测试,删掉后回复更早开始且质量无明显下降;中途想起的事作为新消息补发而不是重启任务(长任务重启代价更高);设计需求列出"不要的风格"清单,比"现代简洁"这类泛泛方向有效得多。

关键要点

  • 一次性交付整个任务 + 完成定义(如"所有 endpoint 迁移完、旧 client 删除、测试全过")+ 唯一停机条件("只在测试因你解释不了的原因失败时来问我")。官方称 Opus 5.5 相比前代最大的提升就在多步长任务:早期测试者让它数小时低监督跑完长编码任务。
  • 删掉 "think carefully / think step by step" 指令:Opus 5.5 每次回复前自动思考并自行决定思考量;官方测试中删除该行后回复更早开始、质量无明显下降。Claude Code 里用 effort 控制思考量,简单问题直接说 "Answer directly"。
  • 长跑控制写进 CLAUDE.md:不需要输入的步骤继续干、状态说明与下一步动作放同一条消息;只在无法继续或破坏性操作前停下(删数据、force-push、改 repo 之外的东西)。破坏性命令的权限提示保持开启。模型经常停在 "Want me to continue?",明确规则能减少这种停车。
  • 大型审计/迁移让模型扇出到 subagent 并逐份核验证据,最后汇总成一张表(service / 是否受影响 / 证据);长任务把 checklist 放进 TASKS.md 文件并随做随更——读文件而非刷屏看进度,文件能挺过上下文压缩对旧轮次的摘要。
  • 结果核查:长跑结束后先读"需要你决策"的部分再读其余总结(可用 CLAUDE.md 固定成 Blocked on me / Changed / Found 三段式);人工 review 前先让模型 review diff——早期测试者称 Opus 5.5 最低 effort 抓的 bug 比 Opus 5 最高 effort 更多且误报更少;研究类任务加一句"标注你未能确认的内容及查过哪里"。
  • Claude apps:图表/截图直接附图、不要手动转打数字(5.5 读图和空间关系理解更强:箭头连接哪些框、两版图改了什么、日历截图里会议几点开始);长文档找错(测试中抓到计划线程里星期对不上的日期、deck 图表与数字不符);要成品文件而非大纲;项目里声明"已回答的内容视为定稿"可避免长对话回溯拖慢响应(长链分析除外,后面的步骤可能推翻前面的)。
  • 误标记处理:Opus 5.5 是首个首发带 Fable 级 bio/cyber 防护的 Opus,被标记的消息大多切到旧模型继续干活;用模型选择器或 /model 切回、Esc×2 编辑重发、/config 改"被标记时切换模型"为先询问、/feedback 反馈误报;检查覆盖整个会话(含文件与搜索结果),所以标记可能来自早前内容。"让它在回复里复现内部推理"本身就是一种触发标记的请求——改问"用三句话解释你为什么选这个方案"。
  • 速度:fast mode 以研究预览形式随 5.5 发布(/fast),同一模型、出字更快,需开启额外用量、每 token 价格高于标准模式。

English Summary

Anthropic's official guide for Opus 5.5 (Addy Osmani, published Sep 22, 2026). Premise: Opus 5.5 works longer on its own, reports plainly on what it did, and decides for itself how much to think before every reply. Recommendations: hand over the whole task in one message with an explicit definition of done and the single reason to stop and ask; delete "think hard/carefully" lines from prompts, project instructions, skills, and CLAUDE.md — in Anthropic's chat-product testing, removing the line made replies start sooner with no clear quality drop; send mid-task reminders as new messages instead of restarting runs (runs are longer now, so restarts cost more); for design work, name the specific styles you don't want rather than "modern and clean". For long Claude Code runs: put a CLAUDE.md rule on when to keep going (status notes in the same message as the next action) versus stop (only when blocked or before destructive actions — deleting data, force-pushing, touching anything outside the repo), keep permission prompts on for destructive commands, fan out audits/migrations to subagents and check each one's evidence into a single table, and keep the task list in a file (TASKS.md) that survives context-window summarization. Checking results: read the "needs from you" part of the report first (format fixed via CLAUDE.md, e.g. Blocked on me / Changed / Found); have the model review diffs before a person does — one early tester said Opus 5.5 at lowest effort caught more bugs than Opus 5 at high effort with fewer false alarms; ask it to mark anything it couldn't confirm and where it looked. In Claude apps: attach charts/screenshots instead of retyping numbers (notably stronger spatial reading), use it to catch contradictions in long documents, ask for the finished file rather than an outline, and mark earlier answers as settled in projects to avoid slow re-processing. Flagged messages: Opus 5.5 ships with Fable-level bio/cyber safeguards and most flagged traffic falls back to an older model — switch back via the model picker or /model, edit and resend with Esc×2, and note that asking the model to reproduce its internal reasoning in the reply is itself a flag category. Fast mode (/fast) is a launch research preview: same model, sooner text, higher per-token cost.

One-liner

Anthropic 官方 Opus 5.5 指南:一次交付整个任务+完成定义、删掉 think-hard 指令、CLAUDE.md 写明何时继续何时停、审计扇出到 subagent 并核验证据、报告先读"需要你决策"的部分。

注:本文件为 content-fetcher cron 回填;opencli 与 web_extract 对 claude.dev 均不可达(TLD 误报/内网判定),最终经浏览器直接抓取 claude.dev 原文全文生成(2026-10-04),含原文发布日期修正(2026-09-22)。