LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
- ID: f09bd61a
- 原文链接: https://arxiv.org/abs/2608.13545
- PDF: https://arxiv.org/pdf/2608.13545v1
- 作者: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
- 日期: 2026-08-13
- 更新: N/A
- 分类: models
- 来源类型: paper
- 标签: curriculum, controlled-corpus, knowledge-acquisition, research-sandbox, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-17T04:23:51+00:00Z
中文导读
LittleLearner:在 LittleCurriculum(88B token仅含美国小学五年级及以下内容显式排除超纲概念/事实/词汇的课程化语料)上从零训练的 5B 模型它语言能力足以支撑开放式评测,但知识与能力边界被课程大纲清晰划界,构成研究知识习得与注入的可控沙盒首批实验考察后训练与上下文学习注入新知识:两种方法都能让模型更好利用已有知识,但不能提升超纲能力语料与模型均开源
为什么值得关注
LittleLearner:在 LittleCurriculum(88B token仅含美国小学五年级及以下内容显式排除超纲概念/事实/词汇的课程化语料)上从零训练的 5B 模型它语言能力足以支撑开放式评测,但知识与能力边界被课程大纲清晰划界。
收录理由:为模型如何习得表示和使用数据提供了边界可解释的受控实验环境,是数据受控研究方法论的范例
关键信息
- 论文标题:LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
- 作者:Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
- arXiv:https://arxiv.org/abs/2608.13545
- 发布时间:2026-08-13
- arXiv 分类:cs.CL, cs.AI, cs.LG
- 关联标签:curriculum, controlled-corpus, knowledge-acquisition, research-sandbox, arxiv
English Abstract
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
English Summary
LITTLELEARNER is a 5B-parameter LLM trained from scratch on LITTLECURRICULUM, a curated 88B-token pretraining corpus restricted to U.S. elementary-school material that explicitly excludes concepts, facts, and vocabulary taught above Grade 5. The model has sufficient language competence for open-ended evaluation while its knowledge and capability boundaries are mapped to interpretable curriculum guidelines, forming a developmentally restricted sandbox for studying knowledge acquisition. First experiments on injecting new knowledge via post-training and in-context learning show both methods help the model better utilize existing knowledge but do not raise out-of-scope capabilities. Corpus and model are both released.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断锚定于条目与论文摘要,未补充摘要之外的内容。