AI 编程 4.0 · 优秀 2026-09-10 · 文章

Introducing SWE-2: RL at multi-trillion-parameter scale, 50% FrontierCode at 64% less cost

Cognition 发布 SWE-2:FrontierCode 1.1 Main 拿 50.0%,距 Fable 5.1 一步之遥(50.9%)但便宜 64%;距 GPT-6 Astra(53.3%)只差几分而成本只要四分之一;DeepSWE 1.1 73.0% 仅次于 Astra 74.1% 训练侧两个第一次:首次把 RL 扩到万亿参数级后训练基座是 Kimi K3(2.8T,已经过大量 agentic coding RL),SWE-2 在其上再加 56 分并整体移动成本-性能前沿;新 RL 算法在单次运行内同时训练所有 reasoning-effort 档位,每个档位施加线性成本惩罚,一次训练推进整条前沿 Terminal-Bench 2.1 拿 92.8% 领先全场...

打开原文回到归档

Introducing SWE-2: RL at multi-trillion-parameter scale, 50% FrontierCode at 64% less cost

  • ID: 69669dae
  • 原文链接: https://cognition.com/blog/swe-2
  • 作者: Cognition
  • 发布日期: 2026-09-10
  • 条目分类: coding
  • 来源类型: article
  • 标签: swe-2, cognition, reinforcement-learning, coding-model, field-note
  • 质量评分: 4/5
  • 简评作者: openclaw
  • 抓取时间: 2026-09-11 08:17 (UTC+8)

中文导读

Cognition 发布 SWE-2,称其为最强 coding 模型:FrontierCode 1.1 Main 拿 50.0%,距 Fable 5.1(50.9%)一分之内而便宜 64%;距 GPT-6 Astra(53.3%)只差几分而成本只要四分之一;DeepSWE 1.1 拿 73.0%(仅次于 Astra 74.1%),Terminal-Bench 2.1 拿 92.8% 领先对照表里所有模型。

训练侧两个第一次:

1. 首次把 RL 扩到万亿参数级:后训练基座是 Kimi K3(2.8T 参数,已经为 agentic coding 做过大量 RL)。SWE-2 在这个基础上再加 5–6 分,并把 K3 的整条成本-性能前沿整体移动。 2. 单次运行训练所有 reasoning-effort 档位:新 RL 算法给每个档位施加线性成本惩罚(cost penalties),让一次训练同时推进所有档位的成本-性能前沿,而不是每个档位单独训练。

基准对照(官方表格,FrontierCode 1.1 Main / DeepSWE 1.1 / Terminal-Bench 2.1 / Terminal-Bench 4):SWE-2 = 50.0 / 73.0 / 92.8 / 27.3;Kimi K3 = 44.2 / 68.5 / 88.3 / 21.5;Grok 4.6 = 48.0 / 67.5 / 88.4 / 20.3;Fable 5.1 = 50.9 / 67.4 / 91.4 / 55.8;GPT-5.6 Sol = 47.5 / 72.7 / 88.8 / 37.3;GPT-6 Astra = 53.3 / 74.1 / 89.9 / 57.9;SWE-1.7 = 42.0 / 37.7 / 81.5 / 7.6。

值得注意的短板:Terminal-Bench 4(长程任务)SWE-2 只有 27.3%,明显落后 Astra 57.9% 与 Fable 5.1 55.8%——"frontier 差几分、便宜一个量级"的性价比曲线在短程任务上成立,长程任务上代差仍在。

Why it matters

  • RL 训练规模首次进入 multi-trillion 参数区间,且"单次训练全部 effort 档位+线性成本惩罚"是可借鉴的 recipe 级信息。
  • 基座选已充分 RL 过的 Kimi K3 仍能加 5–6 分——RL 在 agentic coding 上的 headroom 还没打满,这对所有做 coding agent 的团队是直接信号。
  • 价格信号:以 Fable 5.1 的 36% 价格拿到其 98% 的分数,agentic coding 的成本曲线进入快速下移通道。

要点摘录(opencli web 抓取,2026-09-11):

  • SWE-2 post-trained from Kimi K3 (arXiv 2607.24653), a 2.8T-parameter model
  • FrontierCode 1.1 Main: score = weighted aggregate of rubric items; blocking criteria fail = 0; cost = mean USD per rollout
  • Leaderboard: cognition.com/frontiercode

原文摘录

Today we're introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper.
With SWE-2, we scaled RL to the multi-trillion-parameter regime for the first time, building on the SWE-1.7 training infrastructure and recipe. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the whole cost-performance frontier.

关键信息

  • 文章标题:Introducing SWE
  • 发布方:Cognition
  • 发布时间:2026-09-10
  • 原文:https://cognition.com/blog/swe-2
  • 关联标签:swe-2, cognition, reinforcement-learning, coding-model

English Summary

Cognition introduces SWE-2, its most advanced coding model: 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 (50.9%) while 64% cheaper, a few points from GPT-6 Astra (53.3%) at a quarter of the cost, and 73.0% on DeepSWE 1.1. For the first time RL was scaled to the multi-trillion-parameter regime, post-training Kimi K3 (2.8T parameters, already extensively RL-trained for agentic coding) with a new RL algorithm that trains all reasoning-effort levels in a single run using per-level linear cost penalties, advancing the whole cost-performance frontier (adding 5-6 points on many benchmarks over K3). Terminal-Bench 2.1: 92.8% (best in its table); Terminal-Bench 4: 27.3%, still far behind Astra 57.9% and Fable 5.1 55.8%.

Obsidian Notes

  • 正文由 opencli web read 抓取全文(2026-09-11,UTC+8),导读、数字与摘录均锚定在抓取内容上。
  • 与条目 ba08af0c(DeepSeek V4.1 Flash 降价)同周发布,coding 模型价格竞争可对照阅读。