Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- 复现 Llama-3.2-3B 成本超 150 万美元SmolLM3-3B 超 70 万美元,预训练因此远离学术界与开源社区Puro-2B 给出一条开源预训练配方:在消费级 RTX 5090 上以 FP8 精度从零训练至多 1.4T token 的 2B 模型集合,最佳模型算力成本低于 6900 美元,评测协议下接近 Qwen2.5-1.5B成本效率来自硬件选择低精度训练hyperball 优化课程式模型平均与数据配方的组合由该集合拟合出 Puro Cost Scaling Law:约 4400 美元(低于 5090 美元)即可达到 Qwen2-1.5B 性能由于掌握完整预训练管线,作者还完成了后训练后预训练数据课程如何影响下游表现的受控案例研究数据代码权重以 Apache 2.0 开源
论文信息
- arXiv ID: 2608.27370
- 作者: Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang 等(11 人)
- 发表: 2026-08-27(更新:2026-08-27)
- 分类: cs.CL, cs.LG
- 备注: 62 pages, 20 figures, 24 tables
- 原文链接: https://arxiv.org/abs/2608.27370
- PDF: https://arxiv.org/pdf/2608.27370v1
- 标签:
pretrainingfp8consumer-gpuscaling-lawopen-source
Abstract(原文)
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection differ in token budgets and selected recipe variants. Our best model is trained at a compute cost of less than \$6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol. This cost efficiency is enabled by a combination of approaches, including hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. Beyond the recipe itself, we provide two additional results. First, across the Puro-2B collection, we derive a Puro Cost Scaling Law that relates training cost to average model performance; the fitted law suggests that about \$4.4K, less than \$5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, as an end-to-end case study, we examine how pretraining data curricula shape downstream performance after post-training. Such controlled studies are enabled by having access to the full pretraining pipeline rather than model weights alone. We release the full training recipe for Puro-2B, including data, code, and model weights under Apache 2.0 at https://huggingface.co/collections/thu-pacman/puro-2b.
核心要点(英文摘要的中文提炼)
- 论文题为 Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090,发表于 arXiv(cs.CL, cs.LG,2026-08-27 提交,2026-08-27 更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
- 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。