模型与实验室 5.0 · 必读 2026-08-31 · X

Qwen3.8-Flash Tech Report:四次手术重做 MoE,激活参数砍到 1/3训练 FLOPs 砍到 1/9

@xiaogaifun(高师傅)8月31日精读 Qwen3.8-Flash Tech Report:125B 主干 + 每 token 激活 6B + 51B n-gram Embedding,激活参数与训练 token 都约为 Qwen3.7-Plus 的 1/3,整体训练 FLOPs 1/9,14 项 Benchmark 8 项超上代剩 6 项最大差 2.6 分四次手术分别是 GDN(替换 Attention 降长上下文计算成本)QSA(深层信息冲淡缓解)Gated Residual(4 条并行 Residual Stream 让模型自动分工出长程通道)n-gram Embedding(Backbone 外 51B 参数...

打开原文回到归档

Qwen3.8-Flash Tech Report:四次手术重做 MoE,激活参数砍到 1/3训练 FLOPs 砍到 1/9

  • ID: 1e61d0ec
  • 原文链接: https://x.com/xiaogaifun/status/2094271716054933824
  • 作者/平台: @xiaogaifun / x
  • 发布日期: 2026-08-31
  • 归档分类: models
  • 标签: qwen3.8、moe、gdn、qsa、muon、scaling-law
  • 质量评分: 5/5
  • 抓取时间: 2026-09-01T23:30+08:00

中文导读

@xiaogaifun(高师傅)8月31日精读 Qwen3.8-Flash Tech Report:125B 主干 + 每 token 激活 6B + 51B n-gram Embedding,激活参数与训练 token 都约为 Qwen3.7-Plus 的 1/3,整体训练 FLOPs 1/9,14 项 Benchmark 8 项超上代剩 6 项最大差 2.6 分四次手术分别是 GDN(替换 Attention 降长上下文计算成本)QSA(深层信息冲淡缓解)Gated Residual(4 条并行 Residual Stream 让模型自动分工出长程通道)n-gram Embedding(Backbone 外 51B 参数,可放 host memory)训练侧把 Muon 做主优化器重拟合 Scaling Law砍掉 Batch Size Warmup(省 18.8% 步骤)报告点名真正的瓶颈是 Evaluation Throughput:需要更便宜的中等规模实验预判架构+post-training 整体效果

为什么值得关注

Qwen3.8-Flash 用 GDN/QSA/Gated Residual/n-gram Embedding 四次手术把 MoE 重做一遍,激活参数砍 1/3训练 FLOPs 砍 1/9

关键信息

  • 文章标题:Qwen3.8-Flash Tech Report:四次手术重做 MoE,激活参数砍到 1/3训练 FLOPs 砍到 1/9
  • 作者/平台:@xiaogaifun / x
  • 原文链接:https://x.com/xiaogaifun/status/2094271716054933824
  • 发布日期:2026-08-31
  • 关联标签:qwen3.8、moe、gdn、qsa、muon、scaling-law

English Summary

An architectural deep-dive on Qwen3.8-Flash: 125B total / 6B active / 51B n-gram embedding, achieving ~1/3 of Qwen3.7-Plus's training compute with 8/14 benchmarks above the previous model and the remaining 6 within 2.6 points. Four surgical changes: GDN (replacing attention to cut long-context cost), QSA (mitigating information dilution in deep layers), Gated Residual (four parallel residual streams auto-dividing into long-range channels), and n-gram embedding (51B params outside the backbone, host-memoryable). Training side uses Muon as primary optimizer, refits scaling law, and removes batch-size warmup to save 18.8% optimizer steps. The team names Evaluation Throughput as the real bottleneck: cheaper mid-scale experiments are needed to predict post-trained behavior.

Obsidian Notes

  • 来源:2026-09-01 AK-RSS Digest(89源精选)/ 每日综合摘要 / 调研 / DeepResearch 视所属主题而定
  • 内容由 opencli 拉取原始来源 + Obsidian 笔记交叉核对生成。
  • 中文导读与价值判断均锚定原文摘要与作者;未补充原文章节之外的细节。
  • 抓取时间戳:2026-09-01T23:30+08:00。