基础设施 4.0 · 优秀 2026-08-20 · 论文

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalizati...

面向检索-free 文档知识内化的三段式后训练框架 IAR:Inject 把源文档转成续写/改写/指令条件重建目标做结构化注入,Align 用仅答案的 QA 监督对齐行为,Recover 与基座指令模型合并以恢复通用能力在 Common Corpus/CCI 语料与 LlamaPhiQwenSmolLM 四个模型家族上,8 个数据-模型设置中 7 个于全部四项指标超过 Vanilla SFT:领域 QA 平均 +3.6pt,通用能力 (IFEval/MMLU/MSBench) 平均 +12.1ptRAG 之外,把固定语料变成参数化知识的可复现路线

打开原文回到归档

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

  • arXiv: 2608.20281
  • Authors: Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou
  • Published: 2026-08-20; categories: cs.CL, cs.AI (21 pages, 4 figures)

Abstract

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.

为什么值得读(AAIF 扫描)

研究「文档知识内化」:把一个固定语料转成模型可用的参数知识,支持推理时不检索的问答(retrieval-free QA)。提出 IAR 三阶段后训练框架,把三件事拆开做:

1. Inject——不做常规继续预训练,而是把源文档转成 continuation、重写、指令条件重建三种目标注入; 2. Align——用「只有答案」的 QA 监督对齐问答行为; 3. Recover——把领域适配模型与基础指令模型合并,恢复通用能力。

实验跨 Common Corpus (CC) 和 CCI 两个语料、Llama/Phi/Qwen/SmolLM 四个模型家族:8 个数据集-模型组合里 7 个在全部四项指标上超过 Vanilla SFT,平均域内 QA 准确率 +3.6 个百分点,通用能力(IFEval/MMLU/MSBench 均值)+12.1 个百分点。LoRA 和 FAPM 能在个别通用指标上赢,但在域内内化同样领先的方案里,IAR 的通用能力画像最强。

对「把企业内部文档烘焙进模型权重」的场景(免检索、低延迟、私有部署)这是一个直接可抄的分阶段配方,正面处理了持续预训练最头疼的灾难性遗忘问题。

Source: https://arxiv.org/abs/2608.20281
Captured: 2026-08-22 (AAIF content-fetcher)