模型与实验室 4.0 · 优秀 2026-08-10 · 论文

Fusion Training for Mathematical Generalization in Large Language Models

面向 Thinking Mode Fusion (TMF) 系统研究思考/非思考模式的训练数据比例与训练 schedule:在数学问题求解基准上发现非对称相互作用增加非思考监督比例会显著降低思考模式准确率,不同 schedule 调节这一权衡,最优 schedule 依赖于数据比例,最终量化出非思考与思考模式监督之间的负相关

打开原文回到归档

Fusion Training for Mathematical Generalization in Large Language Models

  • ID: f0b35542
  • Original: https://arxiv.org/abs/2608.09893
  • PDF: https://arxiv.org/pdf/2608.09893v1
  • Authors: Congfeng Cao, Pengyu Zhang, Jelke Bloem
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv Categories: cs.CL, cs.AI
  • AAIF Category: models
  • Source Type: paper
  • Tags: training, reasoning, math, thinking-mode, fusion
  • Quality Score: 4/5
  • Fetched At: 2026-08-12T04:29:18.480369+00:00

Chinese Guide

面向 Thinking Mode Fusion (TMF) 系统研究思考/非思考模式的训练数据比例与训练 schedule:在数学问题求解基准上发现非对称相互作用增加非思考监督比例会显著降低思考模式准确率,不同 schedule 调节这一权衡,最优 schedule 依赖于数据比例,最终量化出非思考与思考模式监督之间的负相关

Why It Matters

面向 Thinking Mode Fusion (TMF) 系统研究思考/非思考模式的训练数据比例与训练 schedule:在数学问题求解基准上发现非对称相互作用增加非思考监督比例会显著降低思考模式准确率,不同 schedule 调节这一权衡...

This content page is grounded in the existing AAIF entry plus arXiv metadata fetched with opencli arxiv paper. It does not add experimental claims beyond the arXiv abstract and entry summary.

Key Metadata

  • Paper title: Fusion Training for Mathematical Generalization in Large Language Models
  • Authors: Congfeng Cao, Pengyu Zhang, Jelke Bloem
  • arXiv: https://arxiv.org/abs/2608.09893
  • PDF: https://arxiv.org/pdf/2608.09893v1
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv categories: cs.CL, cs.AI
  • Comment: ACL SRW 2026
  • Related AAIF tags: training, reasoning, math, thinking-mode, fusion

English Abstract

Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two modes, remain underexplored. In this work, we present a systematic study of TMF by analyzing the effects of the training schedule and data ratio between thinking and non-thinking modes. Focusing on mathematical problem solving, we construct a benchmark with multiple thinking-to-non-thinking data ratios and three training schedules. Our results reveal an asymmetric interaction between the two modes: increasing the ratio of non-thinking supervision reduces the accuracy of the thinking mode. We further show that different training schedules modulate this trade-off and that the optimal schedule depends on the data ratio. Finally, we quantify a negative correlation between non-thinking and thinking mode supervision, highlighting an inherent tension between these two modes. These findings provide practical guidance for designing effective TMF training settings. All code and data are released to support further research at: \href{https://github.com/caocongfeng/Fusion-Bench.git}{\textbf{Fusion Bench}}.

English Summary

Systematic study of Thinking Mode Fusion (TMF) unifying a non-thinking mode and a thinking mode within a single LLM by varying the data ratio and training schedule between the two modes on a math reasoning benchmark. Results show an asymmetric interaction: increasing non-thinking supervision reduces thinking-mode accuracy; different schedules modulate this trade-off; and the optimal schedule depends on the ratio. Quantifies a negative correlation between non-thinking and thinking-mode supervision, framing it as an inherent tension with practical guidance for TMF training settings. Accepted at ACL SRW 2026.

Obsidian Notes

  • Generated by the AAIF content-fetcher cron mode from arXiv metadata and the existing AAIF entry.
  • Chinese guide and relevance notes are anchored to the entry summary, arXiv abstract, authors, dates, categories, and links.
  • Canonical content path is content/{entry_id}.md under the repository root; openclaw/content/ is intentionally avoided.