模型与实验室 4.0 · 优秀 2026-08-27 · 论文

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

RLVR 通常按能力训练领域专家再合并,本文按复用工件划分三种融合范式:Merge 合并专家任务向量,Mix RL 汇总数据集训练,MOPD(多教师在线蒸馏)两者兼用在共享专家与数据跨模型规模与多领域基准下的系统对比显示:三者平均分差最多 1.4 分,但单一基准上差距可达 8.6 分,领域级波动与任务向量几何中的跨领域关系对应训练动态揭示各自约束:Mix RL 受域配比支配,MOPD 受教师上限约束,Merge 把所有专家更新压进一个向量三者都提升单样本准确率,但没有可测的解空间覆盖增益,held-out 能力也无损失实践指南:已有专家且要求廉价融合用 Merge;无专家从头训统一模型用 Mix RL 并调整域配比促进跨域迁移;更看重保留领域增益时用 MOPD

打开原文回到归档

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
  • RLVR 通常按能力训练领域专家再合并,本文按复用工件划分三种融合范式:Merge 合并专家任务向量,Mix RL 汇总数据集训练,MOPD(多教师在线蒸馏)两者兼用在共享专家与数据跨模型规模与多领域基准下的系统对比显示:三者平均分差最多 1.4 分,但单一基准上差距可达 8.6 分,领域级波动与任务向量几何中的跨领域关系对应训练动态揭示各自约束:Mix RL 受域配比支配,MOPD 受教师上限约束,Merge 把所有专家更新压进一个向量三者都提升单样本准确率,但没有可测的解空间覆盖增益,held-out 能力也无损失实践指南:已有专家且要求廉价融合用 Merge;无专家从头训统一模型用 Mix RL 并调整域配比促进跨域迁移;更看重保留领域增益时用 MOPD

论文信息

Abstract(原文)

Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely been studied in isolation, how they compare and how to choose among them remain unclear. We compare all three using shared experts and data across model scales and a multi-domain benchmark suite. Although their average performance differs by at most 1.4 points, the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-domain relations visible in task-vector geometry. Training dynamics expose distinct constraints: Mix RL depends on domain mixture proportions, MOPD remains bounded by its teachers, and Merge compresses all expert updates into one. All three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. These results yield a practical guideline: use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost.

核心要点(英文摘要的中文提炼)

  • 论文题为 Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms,发表于 arXiv(cs.CL,2026-08-27 提交,2026-08-27 更新)。
  • 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
  • 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。