模型与实验室 5.0 · 必读 2026-08-14 · 文章

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Z.ai 发布 GLM-5.3:与 GLM-5.2 同一基座,全部提升来自 post-training 环境扩展Z.ai Code Bench 较 5.2 提升 50%,Terminal Bench 3.0 从 4.6 升至 28.3,DeepSWE v1.1 达 66.9;安全方向 CyberGym 84.5 为 SOTA,ExploitBench 54.4 较 5.2 翻倍权重将在发布两周后开源收录理由:前沿开源权重模型在编码与攻防两条线同时刷新,其环境合成管线(research agent 从真实工作采集任务模式生成长程环境,judge agent 验证可解性)是可借鉴的 post-training 工程方法

打开原文回到归档

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

  • ID: e6c53596
  • 原文链接: https://z.ai/blog/glm-5.3
  • 作者: Z.ai(Research)
  • 日期: 2026-08-14
  • 分类: models
  • 来源类型: blog
  • 标签: glm-5.3, zai, post-training, coding-model, cyber-capability, open-weights
  • 质量评分: 5/5
  • 抓取时间: 2026-08-15T04:45:00Z

中文导读

Z.ai 发布 GLM-5.3:与 GLM-5.2 同一基座模型,全部提升来自 post-training 环境扩展(more environments, more diverse tasks, more compute)。编码能力上,自研 Z.ai Code Bench 较 5.2 提升 50%,Terminal Bench 3.0 从 4.6 升至 28.3(开源 SOTA),DeepSWE v1.1 从 46.2 升至 66.9,Agents' Last Exam 从 23.8 升至 28.5。安全方向涌现能力超出预期:CyberGym 漏洞发现 84.5 为全场 SOTA(超 GPT-5.6 Sol 的 83.6 与 Mythos 5 的 83.8),ExploitBench 54.4 较 5.2 的 24.4 翻倍;且越靠近利用链上游,提升幅度越大。真实世界验证:与多家安全团队合作,在 269 个项目中识别 2,436 个漏洞(1,097 个中高危),最老的漏洞可追溯至 1981 年,平均潜伏 26.6 年;公开披露台账见 cvd.z.ai。权重将在发布两周、安全加固完成后开源。API 变化:thinking 不可再关闭,reasoning_effort 支持 low/high/max(默认 max)。

为什么值得关注

前沿开源权重模型在编码与攻防两条线同时刷新。对工程侧最有参考价值的是其环境合成管线:research agent 从真实工作中采集任务模式,生成长程、多步依赖、含隐藏状态的可运行环境;judge agent 在不接触参考解的情况下试解验证可解性,通过 oracle / no-op / unsolved-state 三类检查后产出可直接训练的二值奖励。这条"环境即数据"的 post-training 路线,与 slime(Megatron 训练 + SGLang rollout、训练- rollout logprob 差控制在 1e-7、端到端 RL 吞吐提升 2.3×)一起,构成可借鉴的工程方法。

关键信息

  • 发布方:Z.ai;日期:2026-08-14;权重两周后开源
  • 基座与 GLM-5.2 相同;提升全部来自 post-training(IndexShare 长上下文、SAO 长程 RL、slime 异步训练栈)
  • 编码基准:Terminal Bench 3.0 28.3(GLM-5.2 为 4.6);DeepSWE v1.1 66.9(46.2);Terminal Bench 2.1 88.2;FrontierSWE 78.1;SWE-Marathon v1.1 42.5
  • 攻防基准:CyberGym 84.5(SOTA);ExploitBench 54.4(24.4 翻倍);ExploitGym 2h/6h 完成 105/130 个任务(5.2 为 29/39)
  • 真实漏洞:269 项目 / 2,436 发现 / 1,097 中高危 / 53 已公开 / 45 年影响跨度(最老 1981)
  • Z.ai Code Bench(私有):Max effort 34.5% @ ~75K output tokens(GLM-5.2 23.4% @ 96K);High effort 31.4% @ ~50K,超 Opus 4.8 的 29.5% @ 120K;仍落后 Fable 5(Max 39.5%)
  • API 迁移:thinking.type: disabled 不再支持,需改为 enabled + reasoning_effort: low|high|max

正文存档(要点摘录)

Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training.

Environment synthesis. The environments now cover production workflows; some represent several days of work for an experienced engineer (e.g., an ML infrastructure task with access to compute clusters, storage, docs, codebases, and experiment results, requiring end-to-end measurable speedup). To scale this, pipelines synthesize environments end to end: research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent attempts each task to verify solvability. Verifiers are synthesized without access to the reference solution; a verifier that passes oracle, no-op, and unsolved-state checks produces a binary reward reliable enough to train on directly.

Emergent cyber capability. As post-training scaled, cyber capability developed faster than expected: the model began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains. Pattern: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2 — and the wider the remaining gap to the closed frontier (Mythos 5: ExploitBench 78.0, ExploitGym 181/247).

Real-world findings. Working with security teams in China: after expert review and dedup, 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity, spanning kernels, OSes, browser engines, OSS infrastructure, web apps, and network protocols; oldest introduced in 1981; average 26.6 years before discovery. Ongoing disclosure tracked in the Z.ai Security Disclosure Ledger (cvd.z.ai): 53 publicly disclosed, 2,383 under embargo.

slime improvements. Top-p mask, top-k/full-vocabulary OPD, R3-style training–rollout consistency setups with full numerical alignment (average logprob difference controlled at 1e-7, >99.99% reduction). Local storage as hierarchical caching layer for multi-teacher OPD; workload-aware heuristics for prefill/decode resource ratios. Net effect: >2.3× end-to-end RL training throughput for long-horizon coding RL.

Performance table (selected):

| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | Qwen3.8-Max | Opus 4.8 | Fable 5 | GPT-5.6 Sol | | --- | --- | --- | --- | --- | --- | --- | --- | | Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 86.6 | 85.0 | 88.0 | 88.8 | | Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | 21.1 | 33.7 | 34.6 | — (col: 34.6 Fable) | | DeepSWE v1.1 | 66.9 | 46.2 | 67.5 | 56.6 | 58.0 | 69.7 | 72.7 | | FrontierSWE | 78.1 | 67.5 | — | — | 66.5 | 88.2 | — | | CyberGym | 84.5 | 77.2 | 80.0 | 78.5 | 78.1 | 83.8 | 83.6 | | ExploitBench | 54.4 | 24.4 | 32.2 | 28.8 | 40.0 | 78.0 | 76.5 | | AutomationBench v1.0.6 | 48.2 | 26.2 | 46.7 | 39.8 | 41.0 | 46.2 | 45.8 | | ALE-CLI | 28.5 | 23.8 | 27.6 | 27.0 | 25.7 | 23.8 | 28.6 |

API changes. Three thinking effort levels: low / high / max (default max, recommended for coding). thinking.type: disabled no longer supported — requests using it will fail; migrate to enabled + reasoning_effort before switching the model ID.

Obsidian Notes

  • 正文存档由 opencli web read 抓取官方博客全文后摘录整理;基准数字、漏洞统计与 API 迁移要求均直接来自原文。
  • 与条目 e6c53596 对应;相关主题:GLM-5.2、slime、SAO、IndexShare、CyberGym。