模型与实验室 4.0 · 优秀 2026-10-01 · 文章

Why do OpenAI's GPT-2 weights beat mine? Part five: data quality

从零复现 GPT-2 系列第五篇:作者在 3.2B token(Chinchilla-optimal)预算下,对 163M 参数模型对比四种数据配方test loss 随 FineWeb 占比线性下降,OpenWebText 最差;IFT 评估(Alpaca 子集微调 + GPT 5.5 当 judge,5 次取平均)上 FineWeb-Edu-only 跑出 24.56,比自家其它模型最高 20.50 高 4 分以上,距 GPT-2 small 的 25.19 只差 0.63他同时发现 curated 数据集训练集与测试集有 13.66% 的 token 重叠,去污染重训后 test loss 几乎不变...

打开原文回到归档

Why do OpenAI's GPT-2 weights beat mine? Part five: data quality

Intake entry · 2026-10-04 · awesome-ai-field-notes

中文摘要

从零复现 GPT-2 系列第五篇:作者在 3.2B token(Chinchilla-optimal)预算下,对 163M 参数模型对比四种数据配方。test loss 随 FineWeb 占比线性下降,OpenWebText 最差;IFT 评估(Alpaca 子集微调 + GPT 5.5 当 judge,5 次取平均)上 FineWeb-Edu-only 跑出 24.56,比自家其它模型最高 20.50 高 4 分以上,距 GPT-2 small 的 25.19 只差 0.63。他同时发现 curated 数据集训练集与测试集有 13.66% 的 token 重叠,去污染重训后 test loss 几乎不变;但同质量的两次 curated 数据选择就能让 IFT 分差 3.05 分——judge 噪声与数据质量信号同量级,所以 FineWeb-Edu 有效只是较强指示而非定论,强结论需要多 seed 重复训练。附:对下一次实验的猜测转向 weight tying(OpenAI 权重预训练时绑定了 embedding 与输出头)。

关键要点

  • 3.2B token 预算下 FineWeb 占比与 test loss 呈线性关系;OpenWebText 配方最差。
  • FineWeb-Edu-only 在 IFT 上 24.56,距 GPT-2 small 的 25.19 仅 0.63,是全系列最接近的一次。
  • curated 训练集 13.66% 测试集 token 污染,去污染后 test loss 几乎不变;但两次同质量数据选择 IFT 分差 3.05,暴露 judge 噪声量级。
  • 下一步杠杆是 weight tying:OpenAI 预训练绑定了 embedding/输出头,而作者的 IFT 测试用的是 163M 非绑定版本。

English Summary

Part five of the GPT-2 reproduction series compares four data recipes at a Chinchilla-optimal 3.2B-token budget for a 163M-parameter model. Test loss falls linearly with FineWeb share (OpenWebText worst); on the IFT eval (Alpaca-subset fine-tune judged by GPT 5.5, averaged over 5 runs) the FineWeb-Edu-only model scored 24.56 - over 4 points above the author's best other model (20.50) and just 0.63 short of GPT-2 small's 25.19. The curated dataset had 13.66% train/test token overlap that barely changed test loss after decontamination, yet two same-quality curated orderings differed by 3.05 IFT points, so judge noise is on the same order as the data-quality signal; strong claims need multi-seed runs. Next lever: weight tying.

One-liner

从零复现 GPT-2 系列第五篇:作者在 3.

原文摘录

The FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got ... and the FineWeb-Edu model is just 0.63 points short of GPT-2 small!
注:本文件为 daily-intake-evening cron 入库;opencli web read 抓取全文后提取,原文全文见抓取缓存。