基础设施 4.0 · 优秀 2026-09-24 · 论文

Paging the Experts: Flash-Backed MoE Inference on iPhone

Routide(arXiv 2609.29032)是一个 Swift/MLX 运行时,在 iPhone 上跑 Qwen3.6-35B-A3B 量化版的文本路径,专家权重放存储,内存只留字节预算内的子集文章刻意保留热节流负面对比和数值等价不成立的限定,做诚实的端侧实测关键数字:五个 128-token 工作负载,固定路由重放下 512 MiB LRU 0.00% demand 命中种子随机驱逐 18.80%576 MiB LRU 38.58%容量悬崖是策略负载交互,不是固定的内存需求同运行时 Mac 对照组保住生成序列在驱逐与异步预取下的一致性,含 2560 次 exact token 比较与 10334 次 speculative load

打开原文回到归档

Paging the Experts: Flash-Backed MoE Inference on iPhone

原文链接: https://arxiv.org/abs/2609.29032
作者: Musa Shams
发布时间: 2026-09-24
源: arxiv

摘要

Routide(arXiv 2609.29032)是一个 Swift/MLX 运行时,在 iPhone 上跑 Qwen3.6-35B-A3B 量化版的文本路径,专家权重放存储,内存只留字节预算内的子集。文章刻意保留热节流、负面对比和「数值等价不成立」的限定,做诚实的端侧实测。关键数字:五个 128-token 工作负载,固定路由重放下 512 MiB LRU 0.00% demand 命中、种子随机驱逐 18.80%、576 MiB LRU 38.58%——容量悬崖是「策略×负载」交互,不是固定的内存需求。同运行时 Mac 对照组保住生成序列在驱逐与异步预取下的一致性,含 2560 次 exact token 比较与 10334 次 speculative load。

English Summary

Routide (arXiv 2609.29032) is a Swift/MLX runtime that runs the text path of a pinned public Qwen3.6-35B-A3B quantized checkpoint on iPhone with expert weights in storage and only a byte-budgeted subset in memory. The paper deliberately preserves thermal throttling, negative results, and 'numerical equivalence does not hold' caveats, doing honest on-device measurements. The headline numbers, across five 128-token workloads with fixed-route replay: 0.00% demand hits with a 512 MiB LRU cache, 18.80% with seeded random eviction at the same budget, and 38.58% with 576 MiB LRU — the apparent capacity cliff is a policy-by-workload interaction, not a universal memory requirement. Same-runtime Mac controls preserve generated sequences across eviction and asynchronous prefetch, including 2,560 exact token comparisons and 10,334 speculative loads.

为什么值得关注

主题线扩展 AAIF moe/flash-storage/iphone/mlx 等主题。

信息源