Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining
Source: https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining
Content fetched: 2026-08-02T23:33:44+08:00
Grounding: scheduled Obsidian digest + local fetched body
Evidence: /Users/gracker/.hermes/evidence/ak-rss-digest/2026-08-02/bodies/15-why-do-openai-s-gpt-2-weights-beat-mine-part-three-testing-overtrainin.txt
Obsidian source: /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/OpenClaw定时任务/AK-RSS-Digest(89源精选)/2026-08-02-AK-RSS-Digest.md
一句话
GPT-2 复现实验通过过训练测试loss 对照和实现排查,展示模型训练差异如何一步步缩小
关键信息
- Author: Giles Thomas
- Original date: 2026-07-31
- Source type: article
- Category: models
- Tags: gpt-2, training-reproduction, overtraining, debugging, model-training
- Quality score: 4
中文摘要
Giles Thomas 的 GPT-2 复现系列保留了真实训练排错过程这一篇围绕 OpenAI GPT-2 权重为何优于个人复现实验,继续测试过训练假设,并用 loss训练轮次数据拆分和实现差异逐步缩小问题范围它的价值不在结论速成,而在展示模型训练复现中如何构造对照实验记录失败路径并定位配置或实现嫌疑
English Summary
Giles Thomas continues a GPT-2 reproduction debugging series by testing whether overtraining explains the gap between his weights and OpenAI's. The post traces loss curves, training duration, data splits, and implementation hypotheses, making it valuable as an example of careful model-training reproduction and experimental debugging.
入库理由
该条来自 2026-08-02 的自动化摘要/书签消化,并已在本地 evidence 中保留原文或长文抓取。它补充 AAIF 中关于模型选择成本、训练复现排错和 coding agent harness 边界的可复用案例。
原文抓取 / Evidence Excerpt
Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining
作者: Giles Thomas
发布时间: 2026-07-31T01:15:00+0000
原文链接: https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining
https://x.com/gpjt https://bsky.app/profile/gilesthomas.com https://github.com/gpjt https://huggingface.co/gpjt https://www.gilesthomas.com/feed/rss.xml
Writing the post that I wished I'd found when I started learning whatever it was...
Archives
Categories
Blogroll
- July 2026 (8)
- June 2026 (7)
- May 2026 (2)
- April 2026 (11)
- March 2026 (3)
- February 2026 (4)
- January 2026 (4)
- December 2025 (1)
- November 2025 (3)
- October 2025 (9)
- September 2025 (3)
- August 2025 (5)
- July 2025 (1)
- June 2025 (2)
- May 2025 (3)
- April 2025 (2)
- March 2025 (7)
- February 2025 (10)
- January 2025 (6)
- December 2024 (7)
- September 2024 (1)
- August 2024 (2)
- July 2024 (2)
- May 2024 (2)
- April 2024 (2)
- February 2024 (2)
- April 2023 (1)
- March 2023 (2)
- September 2022 (1)
- February 2022 (1)
- November 2021 (1)
- March 2021 (1)
- February 2021 (2)
- August 2019 (1)
- November 2018 (1)
- May 2017 (1)
- December 2016 (1)
- April 2016 (1)
- August 2015 (1)
- December 2014 (1)
- August 2014 (1)
- March 2014 (1)
- December 2013 (1)
- October 2013 (3)
- September 2013 (4)
- August 2013 (2)
- July 2013 (1)
- June 2013 (1)
- February 2013 (1)
- October 2012 (1)
- June 2012 (1)
- May 2012 (1)
- April 2012 (1)
- February 2012 (1)
- October 2011 (1)
- June 2011 (1)
- May 2011 (1)
- April 2011 (1)
- March 2011 (1)
- February 2011 (1)
- January 2011 (1)
- December 2010 (3)
- November 2010 (1)
- October 2010 (1)
- September 2010 (1)
- August 2010 (1)
- July 2010 (1)
- May 2010 (3)
- April 2010 (1)
- March 2010 (2)
- February 2010 (3)
- January 2010 (4)
- December 2009 (2)
- November 2009 (5)
- October 2009 (2)
- September 2009 (2)
- August 2009 (3)
- July 2009 (1)
- May 2009 (1)
- April 2009 (1)
- March 2009 (5)
- February 2009 (5)
- January 2009 (5)
- December 2008 (3)
- November 2008 (7)
- October 2008 (4)
- September 2008 (2)
- August 2008 (1)
- July 2008 (1)
- June 2008 (1)
- May 2008 (1)
- April 2008 (1)
- January 2008 (4)
- December 2007 (3)
- March 2007 (3)
- February 2007 (1)
- January 2007 (2)
- December 2006 (4)
- November 2006 (18)
- AI (95)
- TIL deep dives (77)
- Python (74)
- LLM from scratch (48)
- Resolver One (34)
- PyTorch (21)
- TIL (21)
- Blogkeeping (19)
- PythonAnywhere (17)
- Linux (16)
- Startups (15)
- Gadgets (13)
- Hugging Face (13)
- NSLU2 offsite backup project (13)
- Funny (11)
- Musings (11)
- Finance (10)
- Fine-tuning LLMs (10)
- C (9)
- JAX (8)
- Personal (8)
- Robotics (8)
- Website design (8)
- 3D (5)
- Quick links (5)
- Rants (5)
- Cryptography (4)
- JavaScript (4)
- Music (4)
- Oddities (4)
- Talks (4)
- Dirigible (3)
- Eee (3)
- GPT-2 mysteries (3)
- Memes (3)
- Politics (3)
- Django (2)
- GPU Computing (2)
- LaTeX (2)
- MathML (2)
- Microprojects (2)
- OLPC XO (2)
- Retro Language Models (2)
- Space (2)
- VoIP (2)
- Copyright (1)
- Golang (1)
- poppy the training box (1)
- Raspberry Pi (1)
- Software development tools (1)
- Agile Abstractions
- antirez
- Astral Codex Ten
- :: (Bloggable a) => a -> IO ()
- David Friedman's Substack
- Econ & Energy
- Entrepreneurial Geekiness
- For some value of "Magic"
- Hackaday
- kaleidic.ai newsletter
- Knowing.NET
- Language Log
- Millennium Hand
- ntoll.org
- Obey the Testing Goat!
- One Useful Thing
- PK
- PythonAnywhere News
- Simon Willison's Weblog
- Societive
- Software Deviser
- Some opinions, held with varying degrees of certainty
- tartley.com
- the singularity is nearer
- Theia Vogel's website
Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining
Posted on 31 July 2026 in GPT-2 mysteries, AI, JAX, Python
The GPT-2-style models that I've been training work really well, and I've even managed to train some that perform better than the original OpenAI small model in terms of cross entropy loss on a test set. But as I wrote previously, there's a mystery: why do they perform worse on my instruction fine-tuning evaluation?
I had various theories about why that might be, and to me, the most plausible-seeming of them was the amount of data they were trained with. As best I can find out, OpenAI's models were, by modern standards, trained on much more data than they should have been, while I'd used the theoretically optimal amount of training data.
To put it in other words, OpenAI's models were overtrained. If I deliberately overtrained my own models, could I match their performance? This post is a write-up of what happened, but so as not to bury the lede -- it didn't seem to help much, if at all.
Let's see why.
Overtraining
Let's start by getting a nice crisp definition of overtraining.
It's important not to confuse overtraining with overfitting. Overfitting is where you train a model so that instead of learning a general rule about the data it's seeing, it learns something very specific to the training data -- for example, this:
...rather than this:
Overfitting is pretty much always a bad thing.
Overtraining, by contrast, is more of a judgement call. For LLMs, it's generally used as a shorthand for "training for more than the Chinchilla-optimal number of tokens". The Chinchilla paper makes a very specific case: if you train a model for roughly 20 times as many tokens as it has parameters, then you'll have as good a model as you can get for that budget in terms of compute. They were arguing against contemporaneous experiments where people were doing things like doubling the number of parameters but training on the same amount of data.
If you overtrain, it means that you're training on more than 20 tokens per parameter. The Chinchilla argument is that instead of doing that, you should scale up the number of parameters and the number of tokens equally to keep the 20x ratio. Because the amount of compute used scales pretty much linearly with both tokens and parameters, you'll spend the same amount and you'll get a better model that way.
Based on that heuristic, if you double your compute budget, then you should scale up your parameter count by 2 and your training token count by the same amount, and by doing that you'll get a better result than you would if you'd naively just doubled the parameters or the tokens, for the same amount of compute time spent.
But overtraining is not always a bad thing. Keeping things Chinchilla-optimal means that you have to keep scaling up the model as you scale up the compute budget, and often you can't do that -- for example, let's imagine you're training a model that's meant to run on mobile phones. You have a hard limit on the number of parameters: what will fit in the target devices' RAM.
And importantly, in general you will still get a better model by overtraining -- just not as much better as you would have done if you had been able to scale up the model as well as the training tokens.
Now, as always, we don't know enough about the original GPT-2 training runs to be sure as to whether they were overtrained, and if so, by how much. But one thing that we do know is that GPT-2 was trained in 2019, three years before the Chinchilla paper came out, so they definitely didn't use it as a heuristic!
One thing that they do say in the GPT-2 paper is that their dataset, WebText, is "a total of 40 GB of text". Assuming 4 bytes per token (a good rule o
<!-- 原文较长,以下省略;完整抓取见 evidence_path。 -->
author was.
But where that saturation point might be -- and indeed how much training you'd need to do to get there -- is not obvious.
It's an annoying place to finish this experiment, but I guess at least an inconclusive result is better than never having run it at all. And it was at least good to see the test loss improvement that I expected.
But given the opportunity cost of tying up poppy in four-day training runs, I think I'll look into other possibilities next. As I was running this experiment, something came up -- and I'll post about that soon. Luckily, this time it won't involve training more base models...
- * *
1. It's also worth noting that by the same, um, token, the extra-large model, with 1,542M parameters, was _undertrained_ because the Chinchilla-optimal number of tokens would have been about 31B. Though that said, see later regarding epochs. ↩
2. You might wonder whether training over the same tokens repeatedly over multiple epochs "counts" for Chinchilla purposes. Is training on 1.6B tokens over two epochs the same as training on 3.2B tokens over one? "Scaling Data-Constrained Language Models" poked into that in 2023, and from the abstract, came to the conclusion that you could do up to four epochs over the same data without losing much value, but after that returns diminished. If that holds for GPT-2, and they really did train for 42 epochs, maybe they wasted a lot of time? I'll need to read that paper in full at some point. ↩
3. There is a table of results later on in this post where you'll be able to compare models easily. ↩
« Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix How I use AI on this blog »