$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
- ID: 071d4a14
- 原文链接: https://arxiv.org/abs/2609.37976
- 作者: Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu
- 日期: 2026-09-29
- 更新: N/A
- 分类: models
- 来源类型: paper
- 标签: reasoning, model-merging, spectral, efficiency
- 质量评分: 3/5
- 抓取时间: 2026-10-01T15:57:56+00:00
中文导读
作者发现推理能力的核心藏在 Thinking 模型权重中、位于对应 Non-thinking 模型主奇异方向所定义投影的零空间里;据此提出 S³——免训练组合一对 Non-thinking/Thinking checkpoint:主子空间内保留 Non-thinking、子空间外换用 Thinking,推理时省 token 而不掉 thinking 训练带来的精度。2B-30B 的 dense 与 MoE、28 个评测环境上 token 开销平均降 27.4%、整体精度升 1.0 个百分点(HMMT25 上 +8.3% 且提速 33%)。机理解释用注意力熵与简化分析模型;数字好看但解释链偏薄,适合当便宜的部署手段先试再信。
论文信息
- arXiv ID: 2609.37976
- 提交日期: 2026-09-29
- arXiv 分类: cs.LG, cs.AI, cs.CL, cs.CV
- 作者: Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu
- 链接: https://arxiv.org/abs/2609.37976
Abstract(arXiv 原文)
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
为什么值得关注
便宜的部署手段:不改训练、只做子空间组合就能拿到 27% token 节省;机理解释偏薄,先试再信。
English Summary
The authors locate the core of reasoning capacity in the Thinking model's weight component inside the null space of the projection defined by the corresponding Non-thinking checkpoint's dominant singular directions. S3 is a training-free composition that keeps the Non-thinking model inside its dominant subspace and swaps in the Thinking checkpoint outside it. Across 2B-30B dense and MoE models over 28 evaluation environments, it cuts token overhead 27.4% on average while improving accuracy by 1.0 point (+8.3% on HMMT25 with 33% token speedup); attention-entropy analysis supports the mechanism.