模型与实验室 5.0 · 必读 2026-09-09 · 论文

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Conce...

Intern-NCP(上海 AI Lab + 上海交大 LUMIA)2026-09-09 在 arXiv 公开 NCP-ArchPreview 技术报告:在 next-token prediction 之外加一条 Next Concept Prediction 通路从隐状态构造 32128 的 product-quantized 概念词表,由独立 Concept Module 预测下一组 4 token 粒度的概念,再回灌到 token decoder;NTP 与 NCP 端到端联合训练,生成接口仍保持 token 自回归模型规模约 8.9B 参数,在 Dolma-3 课程上训练约 5.73T tokens论文报告:消耗约 51.3% 的总训练 token 量即达到 OLMo-3-7B 最终预训练 loss...

打开原文回到归档

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

Abstract (opencli arxiv paper)

arXiv 2609.10715 abstract (opencli arxiv paper 2609.10715 -f json): 8.9B parameters, trained on 5.73T tokens from Dolma-3; NTP + Next Concept Prediction with product-quantized concept vocabulary; achieves OLMo-3-7B final pretraining loss with 51.3% of training tokens; downstream macro-average +2.45, GSM8K +5.99; 85% of standard compute approaches parameter-aligned 8.9B baseline; 17M-param VQ module for domain adaptation; concept-injected DFlash2 drafter improves mean accepted length by 4.17%.

论文要点 (中文)

Intern-NCP(上海 AI Lab + 上海交大 LUMIA)2026-09-09 在 arXiv 公开 NCP-ArchPreview 技术报告:在 next-token prediction 之外加一条 Next Concept Prediction 通路——从隐状态构造 32×128 的 product-quantized 概念词表,由独立 Concept Module 预测下一组 4 token 粒度的概念,再回灌到 token decoder;NTP 与 NCP 端到端联合训练,生成接口仍保持 token 自回归。模型规模约 8.9B 参数,在 Dolma-3 课程上训练约 5.73T tokens。论文报告:消耗约 51.3% 的总训练 token 量即达到 OLMo-3-7B 最终预训练 loss;全量预训练后下游 macro-average +2.45、GSM8K +5.99;用约 85% 的标准算力逼近同参 8.9B baseline 的训练 loss。概念层还能二次利用:只更新 17M 参数的 VQ 模块可做轻量 domain adaptation;把概念表示注入 DFlash2 draft 模型使 mean accepted length +4.17%(HumanEval +7.59%)。权重以 Apache-2.0 在 Hugging Face ArchSpace-Collection 开源,加载需 trust_remote_code=True。训练代码仍写 coming soon,评测仓在 LUMIA-Group/ncp_olmo_eval。

Key claims (English)

NCP-ArchPreview (Intern-NCP, Shanghai AI Lab + SJTU LUMIA, arXiv 2609.10715, 2026-09-09) adds a Next Concept Prediction objective alongside standard NTP: a product-quantized concept vocabulary (32 codebooks × 128 codewords) is built from hidden states, an 8-layer Concept Module predicts the next 4-token-span concept, and predicted concepts are fed back into the 16-layer token decoder. NTP and NCP are jointly trained end-to-end; generation remains token-level autoregressive. The 8.9B-parameter model was trained on 5.73T Dolma-3 tokens. The paper reports: 51.3% of training tokens reach the OLMo-3-7B final pretraining loss; downstream macro-average +2.45 and GSM8K +5.99; 85% of standard compute approaches the parameter-aligned 8.9B baseline. After pretraining, the latent space is reused: only the 17M-param VQ module is updated for lightweight domain adaptation, and injecting concept representations into a DFlash2 drafter improves mean accepted length by 4.17% (HumanEval MAL +7.59%). Weights are released as Apache-2.0 under Hugging Face ArchSpace-Collection (NCPOlmo3ForCausalLM, trust_remote_code=True); training code is still 'coming soon', evaluation code at LUMIA-Group/ncp_olmo_eval.

病毒帖 vs 论文原句对照

@HowToPrompt__ 在 2026-09-13 发布的解读帖(约 4919 likes / 29.7 万 views)把叙事写成 "ends current LLM era / thinks in concepts instead of predicting one word at a time"——abstract 实际写的是 "preserving standard token-level autoregressive generation",机制是 NTP+NCP 联合 而非替代。"51.3% training data" 指的是 token 预算 达到 OLMo-3-7B 最终预训练 loss 的比例,"half tokens → full convergence" 同样应读为 token-budget 收敛;Stage-1 模型卡明写这个比较 "does not measure wall-clock… or inference throughput"。"massive 6-point math/reasoning" 最贴近的是 GSM8K +5.99(39.27→45.26),Overall AVG 是 +2.45。

Obsidian 证据摘录

ing-NCP-ArchPreview-病毒帖与论文分层-研究材料/master-research]]

NCP-ArchPreview 在 X 上被写成「结束 next-token 时代」:论文原句写了什么

2026 年 9 月 13 日,@HowToPrompt__ 一条解读帖把上海 AI Lab 与上海交大 LUMIA 的 NCP-ArchPreview 推到约 4919 likes、29.7 万 views(fxtwitter 核数,会随时间变)。帖子开场就是:这篇论文「might end the current LLM era」;模型「thinks in concepts instead of predicting one word at a time」。

论文本身(arXiv:2609.10715,2026-09-09 提交)和 Hugging Face 模型卡写的是另一

链接