Finetuning with Sampling: SFT Learns Better Than You Think
- ID: a141568b
- 原文链接: https://arxiv.org/abs/2610.02140
- PDF: https://arxiv.org/pdf/2610.02140
- 作者: Aayush Karan, Sitan Chen, Yilun Du
- 日期: 2026-10-01
- 更新: 2026-10-01
- 分类: models
- 来源类型: paper
- 标签: sft, reinforcement-learning, post-training, mcmc, generalization
- 质量评分: 4/5
- Fetch: 2026-10-03T04:21:07Z
中文导读
传统观念认为 RL 泛化强且不伤已有能力,SFT 泛化弱且易灾难性遗忘;但 SFT 能用 off-policy 专家数据,RL 依赖模型自己重复采样找到成功轨迹本文不改学习目标,而是改数据分布:提出一个 MCMC 采样算法,借助参考模型把 off-policy 轨迹逐步变换得更 on-policy在科学技能习得数学推理开放域专业知识等任务上,该采样让 SFT 与主流 posttraining 技术打平,常常泛化更好遗忘更少,且微调后的模型有较强的分布性表现,能学到超越 sharpening 基座的内容
为什么值得关注
不改目标函数改数据分布:MCMC 把 off-policy 轨迹变 on-policy,让 SFT 打平 RL 且遗忘更少
以上导读与价值判断锚定论文摘要与元数据,完整英文摘要见下文。
关键信息
- 论文标题:Finetuning with Sampling: SFT Learns Better Than You Think
- 作者:Aayush Karan, Sitan Chen, Yilun Du
- arXiv:https://arxiv.org/abs/2610.02140
- 发布时间:2026-10-01
- arXiv 分类:cs.LG, cs.AI, cs.CL
- 关联标签:sft, reinforcement-learning, post-training, mcmc, generalization
English Abstract
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
English Summary
Conventional wisdom: RL generalizes without losing capabilities, SFT generalizes poorly and forgets catastrophically - but SFT learns from off-policy expert data while RL must find successful trajectories by repeated sampling. Rather than modifying the learning objective, the paper tailors the data distribution: an MCMC sampling algorithm progressively transforms off-policy traces to be more on-policy given a reference model. Across scientific skill acquisition, mathematical reasoning, and open-ended expertise, sampling-enabled SFT rivals prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines, with strong distributional performance beyond sharpening the base model.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。