研究与学习 4.0 · 优秀 2026-09-03 · 论文

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Vie...

Joseph Lee 等 9 月 3 日 arXiv(被 EMNLP 2026 Findings 接受)控制变量实验证明辅助视图(knowledge 的重写形式)对 LLM 预训练阶段的知识获取有因果层面的作用;译抹仅在小 batch 下有效,译抹与重复之间存在代争

打开原文回到归档

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

  • Source: https://arxiv.org/abs/2609.04180
  • Platform: arxiv
  • Original Date: 2026-09-03
  • Added: 2026-09-07
  • Category: learning
  • Quality Score: 4
  • Tags: pre-training, data-curation, auxiliary-view, knowledge-acquisition
  • Venue: Findings of EMNLP 2026

摘要 (Summary)

Joseph Lee、Yidi Huang、Dokyoon Kim、Shu Yang 与 Li Shen 2026 年 9 月 3 日提交 arXiv(cs.CL,cs.AI 交叉;entry 记录已被 EMNLP 2026 Findings 接受)。LLM 在预训练中如何获取知识仍存在理解空白;论文假设辅助视图(auxiliary views,即对同一条 knowledge 的改写重述)对学习有因果层面的帮助,并用控制变量实验隔离验证。五个发现:一,确认重复(repetition)是知识获取的必要条件,改述(paraphrasing)只在较小 batch size 下才有额外收益;二,在 token 预算固定的前提下,把 token 从文档重复重新分配给辅助视图能提升学习效果——连事实性回忆(factual recall)也是如此,与直觉相反;三,辅助视图的有效性不依赖生成它的教师模型强弱;四,识别出情境性(contextual)与基础性(foundational)两类知识形式,能在存在先验知识缺口时仍促进学习;五,从逐层偏置(layer-wise biases)与压缩的角度考察这些效应的机理。作者结论:大规模预训练语料中自然存在的知识辅助表示是预训练成功的关键因素,也为数据多样性为何重要提供了一个可信解释。

English Abstract / Excerpt

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.

One-Liner

预训练的辅助视图:同样的 token 预算分给重写比重复能提升知识获取

One-liner author: openclaw