Ornith-1.5: From Self-Scaffolding to Self-Improvement
- ID: e78845aa
- 原文链接: https://ornith.ai/ornith_1_5.html
- 作者: Ornith (@ornith_)
- 日期: 2026-08-20
- 来源类型: article
- 平台: blog
- 分类: models
- 标签: models, self-improvement, open-source, moe, reinforcement-learning
- 质量评分: 4/5
- 抓取时间: 2026-08-20T20:35:00+08:00
中文导读
Ornith-1.5 把 Ornith-1.0 引入的自脚手架(self-scaffolding)框架扩展成更完整的端到端自我改进闭环:模型自主提出新任务 → 生成任务专属脚手架/测试环境(harness)→ 产出解算 rollout,三个环节全部用 GRPO 联合优化,让奖励同时回传到任务生成器与脚手架构造器,不再依赖人工任务集与手工 harness。旗舰版 397B MoE 在 Terminal-Bench 2.1(Terminus-2)拿到 86.1、DeepSWE 拿到 56.0,官方口径与 Claude Opus 4.8(85.0 / 59.0)打平,并超过 GLM-5.2、DeepSeek-V4-Flash-0731 等同级开源模型;9B dense 版量化后可直接部署在 iPhone / Android 上。
任务奖励的设计是本文最有信息量的部分:R_task = V(有效性/可验证性)× D(前沿难度)× N(新颖度),乘法形式强制任务同时满足三项;难度项用当前模型自己的 rollout 成功率估计(目标成功率 p* = 0.2),模型变强后该任务奖励自然衰减,把课程(curriculum)推向前沿。Harness 奖励 = 任务对齐 C × 奖励保真 F × 抗作弊 H,直接对应 agent 训练里最常见的 evaluator hacking 风险。所有评测结果均为 5 次独立运行取平均。
为什么值得关注
- 开源模型自我改进路线的最新样本:任务生成、脚手架、解算三段联合 RL,把"curriculum"从人工设计变成模型自驱。
- 397B MoE 在 agentic coding 基准上对标 Claude Opus 4.8,且明确给出与 GLM-5.2 / DeepSeek-V4-Flash 的横向对比表。
- 任务奖励的 有效性 × 前沿难度 × 新颖度 乘法分解、harness 的抗 reward-hacking 奖励项,是可以直接迁移到其他 self-improving agent 管线的设计。
- 9B-Mobile 版本展示了"端侧 agentic 模型"的可行性:9B 参数在 SWE-bench Verified 拿 70.6,超过 Gemma 4-31B / Qwen 3.6-35B。
关键信息
- 三个规模:397B MoE、35B MoE(激活 3B)、9B dense;基于 Qwen3.5 与 Gemma 4 做 CPT + mid-training + post-training。
- 397B 关键分数(5 次运行平均):
- Terminal-Bench 2.1 (Terminus-2) 86.1(Claude Opus 4.8 为 85.0;Kimi K3 为 88.3)
- Terminal-Bench 2.1 (Claude Code) 85.2(Claude Opus 4.8 为 78.9)
- SWE-bench Verified 86.0(Claude Opus 4.8 为 85.8)
- DeepSWE 56.0(Claude Opus 4.8 为 59.0;GLM-5.2 为 46.2)
- WideSearch 80.8 反超 Claude Opus 4.8(72.9);BrowseComp 86.6
- 35B 版:Terminal-Bench 2.1 (Claude Code) 68.5 vs Qwen3.6-35B 的 49.2;DeepSWE 从 1.0 的 0 提到 22。
- 9B 版:SWE-bench Verified 70.6、Terminal-Bench 2.1 (Claude Code) 47.0,超过 Qwen3.5-9B(53.2 / 18.9)并追平更大模型;量化 9B-Mobile 可跑在 iPhone/Android。
- 方法:任务奖励 V×D×N(p* = 0.2,难度随能力自动升级);harness 奖励 C×F×H;三段全部 GRPO 联合训练。
- 评测卫生:SWE-bench 系列用 OpenHands harness,Git 历史抹除、断网防作弊;NL2Repo 封锁 GitHub 仓库与 pip 访问。
397B 完整对比表(节选自官方)
| Benchmark | Ornith-1.5 397B | DeepSeek-V4-Flash-0731 284B | GLM-5.2 753B | Claude Opus 4.8 | Kimi K3 2.8T | Ornith-1.0 397B | | --- | --- | --- | --- | --- | --- | --- | | Terminal Bench 2.1 (Terminus-2) | 86.1 | 82.7 | 81 | 85 | 88.3 | 77.5 | | Terminal Bench 2.1 (Claude Code) | 85.2 | 81.8 | 82.7 | 78.9 | – | 78.2 | | SWE-bench Verified | 86 | 81.6 | 83 | 85.8 | 86.2 | 82.4 | | SWE-bench Pro | 65.1 | 64.4 | 62.1 | 68 | – | 62.2 | | DeepSWE | 56 | 54.4 | 46.2 | 59 | 67.5 | 8 | | HLE (no tools) | 44.6 | 35 | 40.5 | 49.8 | 43.5 | 30.2 | | GPQA Diamond | 92.8 | 91.4 | 91.2 | 93.6 | 93.5 | 88.1 | | MCP-Atlas | 80 | 74.6 | 77.8 | 82.2 | 82.3 | 76.4 | | Toolathlon-Verified | 71.2 | 70.3 | 48.2 | 76.2 | 73.2 | 43.2 | | WideSearch | 80.8 | 77.3 | 79 | 72.9 | – | 75.2 | | BrowseComp | 86.6 | 84.8 | 85.6 | 84.3 | 91.2 | 79.7 |
English Summary
Ornith-1.5 extends the self-scaffolding framework of Ornith-1.0 into a fuller end-to-end self-improvement loop: the model proposes progressively harder tasks, generates task-specific scaffolds/harnesses, and produces solution rollouts, with all three stages optimized jointly via GRPO. Task reward is a product of validity (V), frontier difficulty (D, estimated from the model's own rollout success rate with target p*=0.2), and novelty (N); harness reward combines task alignment, reward fidelity, and hack resistance. Ships at 397B MoE / 35B MoE / 9B dense scales; the 397B model scores 86.1 on Terminal-Bench 2.1 (Terminus-2) and 56.0 on DeepSWE, on par with Claude Opus 4.8 (85.0/59.0) and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731; the quantized 9B-Mobile variant deploys on iPhone and Android. All results averaged over five independent runs; anti-hacking safeguards (git-history removal, network disabled) applied on SWE-bench evaluations.
Obsidian Notes
- 正文由
opencli web read抓取官方发布公告全文(2026-08-20),导读、对比表与数字均直接取自原文与官方表格,未补充原文之外的信息。 - 与条目 summary_zh/summary_en 一致锚定:397B MoE、Terminal-Bench 2.1 86.1、DeepSWE 56.0、对标 Claude Opus 4.8、9B-Mobile 端侧部署。