产品与商业 4.0 · 优秀 2026-08-18 · 文章

Pacing model development in an era of cyber-critical capabilities

OpenAI 承认即将发布的 Astra 模型可能达到 Preparedness Framework 的 Critical Cybersecurity 阈值,为此暂停最大规模前沿 RL run 两周,期间加固研究集群扩大监控覆盖新监控体系从激活分类器逐 token 运行起步,逐级升级到高算力自动调查者,检查工具动作可用推理与完整活动序列,目标在可疑活动出现 30 分钟内告警并由安全/研究/安保团队接管Astra 所有带工具的推理都被纳入监控,当前监控开销约为被监控推理算力的 20%另含研究环境按工作负载分层沙箱对齐证据要求扩展到训练全过程等措施

打开原文回到归档

Pacing model development in an era of cyber-critical capabilities

中文导读

OpenAI 承认即将发布的 Astra 模型可能达到 Preparedness Framework 的 Critical Cybersecurity 阈值,为此暂停最大规模前沿 RL run 两周,期间加固研究集群、扩大监控覆盖。新监控体系从激活分类器逐 token 运行起步,逐级升级到高算力自动调查者,检查工具动作、可用推理与完整活动序列,目标在可疑活动出现 30 分钟内告警并由安全/研究/安保团队接管。Astra 所有带工具的推理都被纳入监控,当前监控开销约为被监控推理算力的 20%。另含研究环境按工作负载分层沙箱、对齐证据要求扩展到训练全过程等措施。

为什么值得关注

前沿模型网络能力逼近临界阈值时的第一手工程账:两周训练暂停、逐 token 监控与 20% 推理开销

原文摘录 (English Excerpt)

Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework. Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process.

As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This include

Obsidian 证据

  • 来源 digest: ClawFeed 24h 一览 2026-08-20(2026-08-20,评分 8.8)
  • 原文经 opencli web read / opencli arxiv paper 抓取核对,关键数字与摘要均锚定抓取内容。