GPT-6 Astra, Looped Transformers, and Hidden Reasoning
- ID: 44084287
- 原文链接: https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and
- 作者: Sebastian Raschka, PhD
- 发布日期: 2026-09-09
- 条目分类: models
- 来源类型: article
- 标签: gpt-6-astra, looped-transformer, recurrent-depth, chain-of-thought, field-note
- 质量评分: 4/5
- 简评作者: openclaw
- 抓取时间: 2026-09-11 08:17 (UTC+8)
中文导读
GPT-6 Astra 发布后有三个热点:性能、The Information 传闻中的 looped transformer / recurrent depth 架构、以及"Astra 在隐藏思维链"的安全担忧。Raschka 从架构视角逐一拆解。
基准印象:Astra 全面超过 GPT-5.6(写作、数学、编码等类别),3D 渲染与动画类 demo 的提升尤其突出;它仍然是 RLVR 训练、产生中间推理轨迹的推理模型。文中引用 NVIDIA CEO 的说法:Astra 用约 10 万块 Grace Blackwell GPU 训练;computer-use 训练流程里 Mac 是环境而非训练机。
环 transformer 是什么:把中间表示反复通过同一组 transformer 块(attention + FFN + norm + shortcut),而不是堆更多块——权重在多趟之间共享。这个思想并不新,可以追溯到 2018 年的 Universal Transformers。
Astra 真的用了环结构吗:仍是未经官方确认的传闻。但 OpenAI 首席科学家说过"现役 frontier 模型的计算图深度在 GPT-4 的 2 倍以内"——这句话既可以用环结构解释,也可能只是普通层数翻倍。作者自己的判断:Astra 的成功主要来自改进的训练配方与训练数据,环结构可能有帮助,但 The Information 高估了它的贡献。
隐藏思维链担忧:OpenAI 从 o1 起就一直对用户隐藏大部分推理轨迹,所以对终端用户没有实质变化;而对模型开发者来说,环 transformer 也不是掩盖思维链的显著因素——思维链是普通文本 token 的 scratch pad,环结构改变的是计算图的深度复用方式,两者不是一回事。
Why it matters
- frontier 模型架构信息不公开时,"recurrent depth"这类传闻无法自查(如果是开源权重,可以直接验证)——这篇给出了判断这类传闻所需的最小知识装备。
- "计算图深度在 GPT-4 的 2 倍以内"是目前唯一的高层官方信号,但它架构中立的措辞本身就是信息:不能当作环结构的确认。
- 对"隐藏思维链"的安全讨论先要分清对象:对用户是产品选择(从 o1 就开始),对开发者是可解释性问题——而环结构与思维链隐藏在机制上是两码事。
要点摘录(opencli web 抓取,2026-09-11):
- 作者立场:looped transformer 不是 Astra 成功的主因,训练配方与数据才是
- 环结构参考资料:Universal Transformers (arXiv 1807.03819)、Nanbeige4.2-3B 等近期工作
- The Information 原始报道:secret-technique-behind-openais-astra-model-sparks-security-concerns
原文摘录
The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. —— OpenAI 首席科学家(引自文中)
In my opinion, the success (i.e., good modeling performance) behind Astra is likely primarily due to other reasons, namely improved training recipes and training data. The looped transformer tweak might help a bit, but I think that The Information is overestimating its contribution.
I don't think that looped transformers are significant contributors towards hiding or obscuring chains of thought.
关键信息
- 文章标题:GPT-6 Astra, Looped Transformers, and Hidden Reasoning
- 作者:Sebastian Raschka, PhD(Ahead of AI)
- 发布时间:2026-09-09
- 原文:https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and
- 关联标签:gpt-6-astra, looped-transformer, recurrent-depth, chain-of-thought
English Summary
Raschka examines three hot topics around GPT-6 Astra: benchmark impressions (a clear leap over GPT-5.6, disproportionately strong at 3D rendering/animation demos; still an RLVR reasoning model; ~100k Grace Blackwell GPUs per NVIDIA's CEO), a looped-transformer primer (intermediate representations pass through the same weight-shared blocks multiple times, an idea dating to Universal Transformers 2018), and the recurrent-depth rumor (no official confirmation; OpenAI's chief scientist says frontier compute-graph depth is within 2x of GPT-4, which could just mean twice the regular blocks). His verdict: Astra's success likely comes primarily from training recipes and data, looped transformers are overcredited by The Information, and looped transformers are not a significant factor in hiding chains of thought — OpenAI has hidden reasoning traces since o1 anyway.
Obsidian Notes
- 正文由
opencli web read抓取全文(2026-09-11,UTC+8),导读、判断与摘录均锚定在抓取内容上。 - 与条目 93682c1d(GPT-6 Astra API 门与价目)、91607928(重写给 Astra 用的 AGENT.md)同属 Astra 发布周观察线,互不重复。