GLM-5.3 is now open-weight
- ID: a51ad146
- 原文链接: https://huggingface.co/zai-org/GLM-5.3
- 发布方: zai-org(智谱 / Z.ai)
- 发布时间: 2026-08-27(模型卡)
- 模型规模: 753B 参数(BF16 / F8_E4M3 / F32)
- 分类: models
- 来源类型: product
- 标签: open-weights, coding, benchmark, release
- 质量评分: 4/5
- 抓取时间: 2026-08-29T12:59:16+08:00
中文导读
智谱/Z.ai 于 2026-08-27 在 Hugging Face 开放 GLM-5.3 权重(zai-org/GLM-5.3,753B 参数,BF16/F8_E4M3/F32)。GLM-5.3 与 GLM-5.2 共用同一底座,全部增益来自后训练:自研 Z.ai Code Bench 提升 50%,Terminal Bench 3.0 取得开源 SOTA(28.3,GLM-5.2 仅 4.6),Agents' Last Exam(ALE-CLI)28.5。模型卡明示后训练规模化带来超预期的涌现网络攻防能力:CyberGym 漏洞发现 SOTA(84.5),利用链基准较 GLM-5.2 翻倍以上(ExploitBench 54.4 vs 24.4,ExploitGym 2h/6h 105/130 vs 29/39)。支持 SGLang / vLLM / Transformers / KTransformers / Unsloth 等部署(含 Ascend NPU 路径),reasoning_effort 支持 low/high/max 三档、默认 max;聊天模板中 clear_thinking 默认 false,聊天场景需显式传 true。
为什么值得关注
开源权重里最强的编程模型之一,且官方主动披露攻防能力跃升——既是本地部署与微调的现实选项,也是安全评估方需要正视的样本;对后训练路线而言,"同底座纯后训练拿到这些增益"本身就是一份公开对照实验。
原文(抓取存档·节选)
> Model card (huggingface.co/zai-org/GLM-5.3, published 2026-08-27)
GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training.
Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:
- Stronger Coding: most capable open-weights model for coding, +50% over GLM-5.2 on
in-house Z.ai Code Bench; open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam.
- Emergent Cyber Capability: cyber capability developed faster than expected as
post-training scaled; SOTA on CyberGym for vulnerability discovery; gains largest
further up the exploitation chain — more than doubles GLM-5.2 on exploitation benchmarks.
Benchmark (GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4 Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol):
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DS-V4 Pro | Qwen3.8-Max | Opus 4.8 | Fable 5 | GPT-5.6 Sol |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | – | – | 21.1 | 33.7 | 34.6 |
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| SWE-Marathon (v1.1) | 42.5 | 19.4 | 48.1 | – | – | 48.8 | 33.1 | 42.5 |
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym (2h / 6h) | 105 / 130 | 29 / 39 | 36 / 70 | – | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | – | 28.8 | 40.0 | 78.0 | 76.5 |
| AutomationBench (v1.0.6) | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents' Last Exam (ALE-CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
Notes:
- reasoning_effort accepts low / high / max; defaults to max. Keep max for leaderboard reproduction.
- clear_thinking defaults to false in the chat template; pass true explicitly for chat.
- Serve locally via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth;
Ascend NPU supported via vLLM-Ascend / xLLM / SGLang.
- Cyber/Exploit evals ran in Claude Code 2.1.207 harness under domain whitelists to
prevent cheating; ExploitGym reports single-run Pass@1 on 869 tasks at 2h/6h budgets.
Obsidian Notes
- 内容由
opencli web read --url https://huggingface.co/zai-org/GLM-5.3 -f md抓取官方模型卡生成;基准表为节选(省略部分行,标记与原文一致)。 - 中文导读与价值判断锚定在模型卡明示的声明(同底座、纯后训练、涌现攻防能力)与公开基准数字上。