模型与实验室 5.0 · 必读 2026-09-30 · 论文

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

FineWeb 质量过滤后,2026 年 6 月网页 token 中 27.5% 被 Pangram 判为 AI 生成,8 月升至 31.1%作者预训练 800 个语言模型系统改变 AI token 与人类 token 配比并拟合 scaling law:数据饥饿的模型加 AI token 初期降低 human text loss,但很快饱和并反转成伤害;在人类文本高预算下 AI token 几乎立刻抬升 loss,而同量新鲜人类 token 持续降 loss...

打开原文回到归档

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

  • ID: 35cf785f
  • 原文链接: https://arxiv.org/abs/2609.40295
  • PDF: https://arxiv.org/pdf/2609.40295v1
  • 作者: Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
  • 发布时间: 2026-09-30
  • 更新: 2026-09-30
  • 分类: models
  • 来源类型: paper
  • 标签: pretraining, scaling-laws, data-quality, synthetic-text
  • 质量评分: 5/5
  • 抓取时间: 2026-10-02T04:30:58Z

中文导读

FineWeb 质量过滤后,2026 年 6 月网页 token 中 27.5% 被 Pangram 判为 AI 生成,8 月升至 31.1%作者预训练 800 个语言模型系统改变 AI token 与人类 token 配比并拟合 scaling law:数据饥饿的模型加 AI token 初期降低 human text loss,但很快饱和并反转成伤害;在人类文本高预算下 AI token 几乎立刻抬升 loss,而同量新鲜人类 token 持续降 loss;Chinchilla 式定律无法预测这一行为新定律用独立的 benefit/harm 两项让 AI token 价值可变号,无 AI 文本时退化为 Chinchilla,外推到 3.6 倍大小模型时误差比现有最佳定律低 41%建议:目标是人类文本时过滤 AI 文本先重复人类文本再扩 AI 语料分开报告人类/AI 验证 loss发布 WildAI 83B token 标注语料全部 800 个模型与代码

为什么值得关注

800 个预训练实验证明野生 AI 生成文本对预训练的收益可变号:数据饥饿时先甜后毒,高预算时几乎只有害

The abstract reports 27.5% of FineWeb-filtered June-2026 web tokens labeled AI-generated by Pangram (31.1% by August); across 800 pretrained LMs, added AI tokens help data-starved runs only briefly before reversing into harm, and the proposed benefit/harm scaling law predicts held-out human-text loss on models up to 3.6x larger with 41% lower error than the best prior law.

关键信息

  • 论文标题:How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
  • 作者:Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
  • arXiv: https://arxiv.org/abs/2609.40295
  • 发布时间:2026-09-30
  • arXiv 分类:cs.CL, cs.LG
  • 关联标签:pretraining, scaling-laws, data-quality, synthetic-text

英文摘要

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.

English Summary

After FineWeb filtering, 27.5% of June 2026 web tokens are labeled AI-generated by Pangram, rising to 31.1% by August. Pretraining 800 LMs across AI:human token ratios shows AI tokens initially lower human-text loss for data-starved models but saturate and reverse into harm, while at high human-data budgets AI tokens raise loss almost immediately; Chinchilla-style laws fail to predict this. A new scaling law with separate benefit and harm terms lets the value of an AI token change sign, reduces to Chinchilla without AI text, and extrapolates to 3.6x larger models with 41% lower error. The authors release the 83B-token labeled WildAI corpus, all 800 models, and code.

Obsidian Notes

  • 本页为内容补齐(content backfill):条目早已入库,本次补写内容页。
  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锺定在条目已有摘要与本次拉取的论文摘要上,未添加摘要之外的实验细节。