GLM-5.3 and the spread of advanced cyber capabilities
- ID: 80e4ef84
- 原文链接: https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- 作者: Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, Tripp Gallagher
- 日期: 2026-09-29
- 分类: industry
- 来源类型: article
- 标签: ai-safety, cyber, red-team, open-weights, ai-policy
- 质量评分: 4/5
- 抓取时间: 2026-09-30T15:48:51Z
中文导读
Anthropic Frontier Red Team 评估 GLM-5.3 与 Claude Mythos Preview 在内部 Binary Exploitation 基准 100 个任务上的表现:GLM-5.3 在 4% 的尝试里完成完整 control flow hijack,Mythos Preview 为 6%,而 Opus 4.6 与 GLM-5.2 全部失败,Anthropic 判断有意义的门槛已被跨过,端到端漏洞利用能力开始扩散模拟测试中攻击者用简单技巧绕过 GLM-5.3 护栏的成功率在 64% 到 100%,受护栏保护的 Claude 模型未被攻破;NIST CAISI 独立评估称 GLM-5.3 是迄今最强 cyber 开源权重模型,与美国前沿差距约四个月Anthropic 的结论是:无护栏开源发布显著提高恶意行为者可用的攻击能力,但同样可被防御者用来加固系统
为什么值得关注
能力扩散证据链的一手来源:开源权重模型首次被头部实验室实测确认具备端到端漏洞利用能力,model release policy 与 cyber 防御假设都要跟着改。
关键信息
- 标题: GLM-5.3 and the spread of advanced cyber capabilities
- 作者: Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, Tripp Gallagher
- 原文链接: https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- 发布时间: 2026-09-29
- 来源平台: anthropic
- 关联标签: ai-safety, cyber, red-team, open-weights, ai-policy
English Summary
Anthropic's Frontier Red Team evaluates GLM-5.3 on 100 random tasks from its internal Binary Exploitation benchmark: GLM-5.3 developed full control flow hijacks in 4% of trials and Claude Mythos Preview in 6%, while earlier models like Claude Opus 4.6 and GLM-5.2 never succeeded a meaningful capability threshold toward proliferated end-to-end cyber exploits. In simulated tests, simple techniques bypass GLM-5.3's safeguards 64% to 100% of the time; safeguarded Claude models held. NIST CAISI independently calls GLM-5.3 the most cyber-capable open-weight model to date, lagging the US frontier by about four months. Anthropic's assessment: releasing capable weights without meaningful safeguards significantly increases capabilities available to malicious actors, while defenders can also benefit.
English Source Excerpt
Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The rapid rate of improvement in AI suggested to us that this ability would eventually proliferate to many other models, making it much easier for malicious cyber actors to launch highly impactful cyberattacks.
In light of these considerations, we chose to release Claude Mythos Preview in a limited way, through Project Glasswing—which enabled trusted cyber defenders to find more than 10,000 vulnerabilities in critical software, giving them a head start before malicious actors had access to similarly capable models.
But those models have now arrived. In this post, we share our analysis of GLM-5.3, the latest AI model developed by Zhipu AI (known outside of China as Z.ai). Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests. In contrast, these attacks did not succeed against safeguarded Claude models in our testing. We assess that GLM-5.3’s lax safeguards significantly increase the cyber capabilities available to malicious ac
Obsidian Notes
- 内容由
opencli web read抓取原文页面生成。 - 中文导读与英文摘要基于当日抓取正文与 Obsidian 摘要笔记交叉核对;英文摘录为原文逐字片段。