Red Alert: OpenAI is poised to cross an AI safety redline
- ID: 90d11ff7
- 原文链接: https://garymarcus.substack.com/p/red-alert-openai-is-poised-to-cross
- 作者: Gary Marcus
- 日期: 2026-09-02
- 抓取时间: 2026-09-02T15:31:00Z
- 分类: industry
- 来源类型: article
- 语言: zh
- 标签: openai, safety, chain-of-thought, monitorability, astra
- 质量评分: 4/5
中文导读
Gary Marcus 9 月 2 日就 The Information 当天爆料做的快评:OpenAI 在 Astra 模型里尝试一种叫 "recurrent depth" 的新推理技术,让模型"思维"过程更接近神经隐喻(neuralese),副作用是直接破坏 chain-of-thought 可读性等于把人类目前监控 LLM 最可靠的那条 thread 砍断Marcus 引用 Chain of Thought Monitorability 论文的观点(CoT 监控是"fragile opportunity"),又贴出前 OpenAI 安全研究员 Steven Adler 的原推"Absolutely do not train your models like this - what is going on??"和 Nathan Calvin 的同步警告Marcus 的判断很硬:CoT 监控本身不完美(Subbarao Kambhampati 早就质疑过),但它是当下唯一对黑盒 LLM 有效的"纤细线索",拿性能的小幅提升去换这条线索是危险赌注
一句话点评
Gary Marcus 9 月 2 日就 The Information 当天爆料做的快评:OpenAI 在 Astra 模型里尝试一种叫 "recurrent depth" 的新推理技术,让模型"思维"过程更接近神经隐喻(neuralese).
English Abstract / Summary
Gary Marcus flags a The Information report that OpenAI is testing "recurrent depth" reasoning in Astra, which pushes model "thinking" closer to neuralese and breaks chain-of-thought readability, severing the most reliable human monitorability thread for black-box LLMs. He cites the Chain of Thought Monitorability paper, Steven Adler's "Absolutely do not train your models like this" reply, and Nathan Calvin's warning, and argues CoT monitoring is imperfect but currently the only thin thread for oversight of black-box models.
Obsidian Notes
- 由
daily-intake-evening2026-09-02 cron 从当日 Obsidian 摘要(论文流水线 / AK-RSS / ClawFeed / X 书签消化)发现并入库存量阶段。 - 中文导读与判断均锚定在条目已有摘要、源页面正文、作者、日期与分类信息;未补充源页面之外的实验细节。