模型与实验室 4.0 · 优秀 2026-09-28 · 产品

Introducing Claude Sonnet 5.5

Anthropic 发布 Claude 5.5 家族第二款模型 Sonnet 5.5:比 Sonnet 5 快 30%+,单价不变($2/$10 每百万输入/输出 token)但单任务成本最多降 30%agentic coding 基准 Terminal-Bench 4.0 得分 70.6%(Sonnet 5 仅 10.3%),GDPval-AA 距 Opus 5.5 仅 2 分,还是首个纯靠截图打通宝可梦红的 Sonnet由于网络安全能力比肩 Opus 5,它成为首个上线 cyber safeguards 与 fallback 的 Sonnet;Haiku 5.5 将在数周内加入家族

打开原文回到归档

Introducing Claude Sonnet 5.5

原文链接: https://www.anthropic.com/claude-sonnet-5-5
作者: Anthropic
发布时间: 2026-09-28
源: arXiv外部扫描 (2026-09-29)

摘要

Anthropic 发布 Claude 5.5 家族第二款模型 Sonnet 5.5:比 Sonnet 5 快 30%+,单价不变($2/$10 每百万输入/输出 token)但单任务成本最多降 30%agentic coding 基准 Terminal-Bench 4.0 得分 70.6%(Sonnet 5 仅 10.3%),GDPval-AA 距 Opus 5.5 仅 2 分,还是首个纯靠截图打通宝可梦红的 Sonnet由于网络安全能力比肩 Opus 5,它成为首个上线 cyber safeguards 与 fallback 的 Sonnet;Haiku 5.5 将在数周内加入家族

English Summary

Anthropic introduces Claude Sonnet 5.5, the second model in the Claude 5.5 family: a clear upgrade over Sonnet 5 that runs 30%+ faster and typically costs up to 30% less per task at the same price ($2/$10 per million input/output tokens). Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 agentic coding (vs 10.3% for Sonnet 5) and two points below Opus 5.5 on GDPval-AA, and is the first Sonnet to beat Pokemon Red working only from screenshots. Because its cybersecurity capabilities are comparable to Opus 5's, it is the first Sonnet model to launch with cyber safeguards and fallbacks like those developed for Anthropic's most capable models; biology safeguards match Sonnet 5's. Haiku 5.5 will join the family in the coming weeks.

为什么值得关注

Sonnet 档首次带上 cyber safeguards:能力比肩 Opus 5 的中档模型开始继承旗舰安全栈,日常 agentic coding 的性价比基准被重设

信息源