GLM-5.3-Flash: Frontier Intelligence, Flash Cost
- ID: 05f53b90
- 原文链接: https://z.ai/blog/glm-5.3-flash
- 作者: Z.ai
- 日期: 2026-08-26
- 分类: models
- 来源类型: article
- 标签: models, glm, hybrid-attention, efficient-inference, agentic-benchmarks
- 质量评分: 4/5
- 抓取时间: 2026-08-27T05:21:12Z
中文导读
智谱 Z.ai 发布 GLM-5 系列首个原生多模态模型 GLM-5.3-Flash:总参数 320B激活参数仅 18B,价格约为前代的十分之一,却在其六个编码与 agentic 基准上全面超过 GLM-5.2(DeepSWE v1.1 得 63.4 对 46.2,AutomationBench 得 48.8 对 26.2),总体接近 Claude Opus 4.8(Z.ai Code Bench 最高努力档 29.0 对 29.5);在 Artificial Analysis Intelligence Index v4.1.1 上以每任务约 0.045 美元的折扣价拿下 57 分架构上首次引入稀疏加线性的混合注意力,显著降低长上下文服务成本,并用 mHC(流形约束超连接)提升缩放效率;1M token 上下文下通过 IndexPool 将四组索引器键向量压缩为一组;对比 GLM-5.3 注意力计算量降为约 1/3.0KV cache 约 1/4.4发布前曾以匿名代号 ox-alpha 在 OpenCode 与 OpenRouter 盲测,迅速登顶当周最受欢迎模型,且全部流量跑在国产 AI 芯片上
为什么值得关注
GLM-5.3-Flash:320B 总参/18B 激活+稀疏线性混合注意力,约 1/10 价格逼近 Claude Opus 4.8;曾以 ox-alpha 匿名盲测登顶当周人气
关键信息
- Scale: 320B total params / 18B active; priced ~1/10 of GLM-5.2; first natively multimodal GLM-5 series model.
- Artificial Analysis Intelligence Index v4.1.1: 57 points at ~$0.045 per task (discounted) - pushes the cost-performance Pareto frontier.
- Coding/agentic vs GLM-5.2: DeepSWE v1.1 63.4 vs 46.2; AutomationBench v1.0.6 48.8 vs 26.2; Toolathlon Verified 78.4 vs 59.9; approaches Claude Opus 4.8 overall (Z.ai Code Bench max effort 29.0 vs 29.5).
- Architecture: hybrid sparse + linear attention; Manifold-Constrained Hyper-Connections (mHC); IndexPool compresses four indexer key vectors into one for 1M-token context; vs GLM-5.3 attention compute /3.0x and KV cache /4.4x.
- Blind-tested anonymously as ox-alpha on OpenCode and OpenRouter; became the most popular model of the week with all traffic served on Chinese AI chips.
- Serving stack: custom engine built on SGLang for domestic accelerators, W8A8 quantization, hybrid INT8/FP8/BF16 cache quantization, Encode-Prefill-Decode (EPD) disaggregated architecture across tens of thousands of chips; 3x end-to-end serving improvement over initial baseline.
- Availability: open weights on HuggingFace (zai-org/GLM-5.3-Flash); SGLang / vLLM / TokenSpeed supported; 3x usable quota for GLM Coding Plan users.
English Summary
Z.ai introduces GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series: 320B total parameters with only 18B active, outperforming GLM-5.2 across six coding and agentic benchmarks (DeepSWE v1.1 63.4 vs 46.2, AutomationBench 48.8 vs 26.2) at roughly one-tenth the price, while approaching Claude Opus 4.8 overall (29.0 vs 29.5 on Z.ai Code Bench at max effort) and scoring 57 on Artificial Analysis Intelligence Index v4.1.1 at ~$0.045 per discounted task. The architecture debuts a hybrid sparse-plus-linear attention design to cut long-context serving cost, adds Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency, uses IndexPool to compress four indexer key vectors into one at 1M-token context, and reduces attention compute and KV cache 3.0x and 4.4x versus GLM-5.3....
Obsidian Notes
- Full article body fetched via
opencli web read --url https://z.ai/blog/glm-5.3-flash --download-images false -f md(2026-08-27T05:21:12Z); archived under the profile workspaceweb-articles/GLM-5.3-Flash_Frontier_Intelligence,_Flash_Cost/. - 中文导读与价值判断锚定在条目已有摘要与论文摘要上,未补充摘要之外的实验细节。
- Canonical backfill page for existing entry
05f53b90indata/entries.json; generated by AAIF content-fetcher (2026-08-27), no push.