模型与实验室 4.0 · 优秀 2026-08-26 · 文章

GLM-5.3-Flash: Frontier Intelligence, Flash Cost

智谱 Z.ai 发布 GLM-5 系列首个原生多模态模型 GLM-5.3-Flash:总参数 320B激活参数仅 18B,价格约为前代的十分之一,却在其六个编码与 agentic 基准上全面超过 GLM-5.2(DeepSWE v1.1 得 63.4 对 46.2,AutomationBench 得 48.8 对 26.2),总体接近 Claude Opus 4.8(Z.ai Code Bench 最高努力档 29.0 对 29.5);在 Artificial Analysis Intelligence Index v4.1.1 上以每任务约 0.045 美元的折扣价拿下 57 分架构上首次引入稀疏加线性的混合注意力,显著降低长上下文服务成本,并用 mHC(流形约束超连接)提升缩放效率...

打开原文回到归档

GLM-5.3-Flash: Frontier Intelligence, Flash Cost

  • ID: 05f53b90
  • 原文链接: https://z.ai/blog/glm-5.3-flash
  • 作者: Z.ai
  • 日期: 2026-08-26
  • 分类: models
  • 来源类型: article
  • 标签: models, glm, hybrid-attention, efficient-inference, agentic-benchmarks
  • 质量评分: 4/5
  • 抓取时间: 2026-08-27T05:21:12Z

中文导读

智谱 Z.ai 发布 GLM-5 系列首个原生多模态模型 GLM-5.3-Flash:总参数 320B激活参数仅 18B,价格约为前代的十分之一,却在其六个编码与 agentic 基准上全面超过 GLM-5.2(DeepSWE v1.1 得 63.4 对 46.2,AutomationBench 得 48.8 对 26.2),总体接近 Claude Opus 4.8(Z.ai Code Bench 最高努力档 29.0 对 29.5);在 Artificial Analysis Intelligence Index v4.1.1 上以每任务约 0.045 美元的折扣价拿下 57 分架构上首次引入稀疏加线性的混合注意力,显著降低长上下文服务成本,并用 mHC(流形约束超连接)提升缩放效率;1M token 上下文下通过 IndexPool 将四组索引器键向量压缩为一组;对比 GLM-5.3 注意力计算量降为约 1/3.0KV cache 约 1/4.4发布前曾以匿名代号 ox-alpha 在 OpenCode 与 OpenRouter 盲测,迅速登顶当周最受欢迎模型,且全部流量跑在国产 AI 芯片上

为什么值得关注

GLM-5.3-Flash:320B 总参/18B 激活+稀疏线性混合注意力,约 1/10 价格逼近 Claude Opus 4.8;曾以 ox-alpha 匿名盲测登顶当周人气

关键信息

  • Scale: 320B total params / 18B active; priced ~1/10 of GLM-5.2; first natively multimodal GLM-5 series model.
  • Artificial Analysis Intelligence Index v4.1.1: 57 points at ~$0.045 per task (discounted) - pushes the cost-performance Pareto frontier.
  • Coding/agentic vs GLM-5.2: DeepSWE v1.1 63.4 vs 46.2; AutomationBench v1.0.6 48.8 vs 26.2; Toolathlon Verified 78.4 vs 59.9; approaches Claude Opus 4.8 overall (Z.ai Code Bench max effort 29.0 vs 29.5).
  • Architecture: hybrid sparse + linear attention; Manifold-Constrained Hyper-Connections (mHC); IndexPool compresses four indexer key vectors into one for 1M-token context; vs GLM-5.3 attention compute /3.0x and KV cache /4.4x.
  • Blind-tested anonymously as ox-alpha on OpenCode and OpenRouter; became the most popular model of the week with all traffic served on Chinese AI chips.
  • Serving stack: custom engine built on SGLang for domestic accelerators, W8A8 quantization, hybrid INT8/FP8/BF16 cache quantization, Encode-Prefill-Decode (EPD) disaggregated architecture across tens of thousands of chips; 3x end-to-end serving improvement over initial baseline.
  • Availability: open weights on HuggingFace (zai-org/GLM-5.3-Flash); SGLang / vLLM / TokenSpeed supported; 3x usable quota for GLM Coding Plan users.

English Summary

Z.ai introduces GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series: 320B total parameters with only 18B active, outperforming GLM-5.2 across six coding and agentic benchmarks (DeepSWE v1.1 63.4 vs 46.2, AutomationBench 48.8 vs 26.2) at roughly one-tenth the price, while approaching Claude Opus 4.8 overall (29.0 vs 29.5 on Z.ai Code Bench at max effort) and scoring 57 on Artificial Analysis Intelligence Index v4.1.1 at ~$0.045 per discounted task. The architecture debuts a hybrid sparse-plus-linear attention design to cut long-context serving cost, adds Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency, uses IndexPool to compress four indexer key vectors into one at 1M-token context, and reduces attention compute and KV cache 3.0x and 4.4x versus GLM-5.3....

Obsidian Notes

  • Full article body fetched via opencli web read --url https://z.ai/blog/glm-5.3-flash --download-images false -f md (2026-08-27T05:21:12Z); archived under the profile workspace web-articles/GLM-5.3-Flash_Frontier_Intelligence,_Flash_Cost/.
  • 中文导读与价值判断锚定在条目已有摘要与论文摘要上,未补充摘要之外的实验细节。
  • Canonical backfill page for existing entry 05f53b90 in data/entries.json; generated by AAIF content-fetcher (2026-08-27), no push.