Scaling Laws for Looped Mixture of Experts
- ID: 5648565c
- 原文链接: https://arxiv.org/abs/2609.40316
- PDF: https://arxiv.org/pdf/2609.40316v1
- 作者: Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
- 发布时间: 2026-09-30
- 更新: 2026-09-30
- 分类: models
- 来源类型: paper
- 标签: scaling-laws, moe, looped-transformer, efficiency
- 质量评分: 4/5
- 抓取时间: 2026-10-02T04:30:58Z
中文导读
首个同时建模循环(looped transformer)与稀疏(MoE)的 scaling law:在模型规模与数据之外引入有界稀疏条件化的 recurrence 映射,刻画 looping 带来的有效参数增益及稀疏性对该增益的放大;对 held-out loss 的预测精度超过既有定律,并把稠密与 MoE 定律作为特例恢复拟合结果给出设计准则:稀疏带来约 3 倍活跃参数效率,recurrence 在推理任务上带来约 2 倍总参数效率,两轴联合进一步推进前沿万亿 token 规模验证:同等训练算力下,按定律选取循环次数的 looped MoE 在推理基准上匹敌约 2 倍大小的非循环 MoE,并通过 recurrence 获得测试时扩展能力
为什么值得关注
循环+稀疏联合 scaling law:稀疏省活跃参数 3 倍循环省总参数 2 倍,万亿 token 规模验证成立
The abstract introduces Loop Scaling Laws that jointly model recurrence and sparsity; sparsity yields ~3x active-parameter efficiency, recurrence ~2x total-parameter efficiency on reasoning, and at matched compute a looped MoE matches a ~2x larger non-looped MoE on reasoning benchmarks at trillion-token scale.
关键信息
- 论文标题:Scaling Laws for Looped Mixture of Experts
- 作者:Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
- arXiv: https://arxiv.org/abs/2609.40316
- 发布时间:2026-09-30
- arXiv 分类:cs.LG, cs.AI, cs.CL
- 关联标签:scaling-laws, moe, looped-transformer, efficiency
英文摘要
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.
English Summary
The first scaling law jointly modeling recurrence (looping) and sparsity (MoE) alongside model size and data, built on a bounded sparsity-conditional recurrence mapping. It predicts held-out loss more accurately than prior alternatives and recovers dense and MoE scaling laws as special cases. Fitted laws imply ~3x active-parameter efficiency from sparsity and ~2x total-parameter efficiency on reasoning from recurrence; at trillion-token scale a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on reasoning benchmarks at matched compute, enabling test-time scaling via recurrence.
Obsidian Notes
- 本页为内容补齐(content backfill):条目早已入库,本次补写内容页。
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锺定在条目已有摘要与本次拉取的论文摘要上,未添加摘要之外的实验细节。