模型与实验室 4.0 · 优秀 2026-09-22 · GitHub

livenerf: deterministic benchmark for post-launch model drift (Opus 5.5)

针对Anthropic 上线后悄悄削弱模型的持续争论,livenerf 从 Opus 5.5 发布日(2026-09-22)起建 day-0 基线:78 题固定题组(2,336 题池中筛出的 GPQA Diamond / MMLU-Pro / competition-math / AIME 2025-26 有时有错的部分),每天 90 个 sample 跑 30 天,再加 10 天 baseline 与两组 10 天对照窗因为采样参数被锁思考无法关闭,项目把其余一切钉死:frozen promptspinned Claude Code CLI 2.1.280harness 哈希 461391b6fce64167永久保留原始日志,基于 UK AISI 的 Inspect 框架...

打开原文回到归档

livenerf: deterministic benchmark for post-launch model drift (Opus 5.5)

  • ID: bf2e738d
  • 原文链接: https://github.com/ninjahawk/livenerf
  • 作者: ninjahawk
  • 日期: 2026-09-22
  • 分类: models
  • 来源类型: github
  • 标签: evaluation、benchmark、opus、model-drift
  • 质量评分: 4/5
  • 抓取时间: 2026-10-02 (daily-intake-evening, opencli/web fetch)

中文导读

针对「Anthropic 上线后悄悄削弱模型」的持续争论,livenerf 从 Opus 5.5 发布日(2026-09-22)起建 day-0 基线:78 题固定题组(2,336 题池中筛出的 GPQA Diamond / MMLU-Pro / competition-math / AIME 2025-26 有时有错的部分),每天 90 个 sample 跑 30 天,再加 10 天 baseline 与两组 10 天对照窗。因为采样参数被锁、思考无法关闭,项目把其余一切钉死:frozen prompts、pinned Claude Code CLI 2.1.280、harness 哈希 461391b6fce64167、永久保留原始日志,基于 UK AISI 的 Inspect 框架,统计沿用 Anthropic 自己的 Adding Error Bars to Evals。validate 阶段已量化 effort 开关的影响:effort low 让输出 token -62%、准确率 -8.3±4.5 分;effort medium 是 -26% token、-4.2±3.9 分;同规模换 Opus 5 在 99% 置信度下与 5.5 不可区分(-3.8±6.3)。截至 2026-10-01 完成 8/30 天,首个 results 行要等 day 20 之后。

为什么值得关注

「模型被悄悄削弱」争论里第一个严肃的公开测量装置:day-0 基线 + 全钉死 harness + Anthropic 自家误差棒方法,已先证明 effort 开关本身就是 -8.3 分。

English Summary

Against the running argument about whether Anthropic quietly degrades models after launch, livenerf builds a day-0 baseline from Opus 5.5's launch (2026-09-22): a frozen 78-question panel (from a 2,336-question pool of GPQA Diamond / MMLU-Pro / competition-math / AIME 2025-26 items the model sometimes misses), 90 samples per day for 30 days, plus a 10-day baseline and two 10-day comparison windows. Since sampling params are locked and thinking can't be turned off, everything else is pinned: frozen prompts, pinned Claude Code CLI 2.1.280, harness hash 461391b6fce64167, raw logs kept forever; built on UK AISI's Inspect with stats following Anthropic's own Adding Error Bars to Evals. Validation already quantifies effort switches: effort low = -62% output tokens and -8.3 +/- 4.5 points; effort medium = -26% tokens and -4.2 +/- 3.9; swapping in Opus 5 was not distinguishable from 5.5 at 99% (-3.8 +/- 6.3). 8 of 30 days collected as of 2026-10-01; first results row after day 20.

Obsidian 原文摘录(抓取正文头部)

# ninjahawk/livenerf: Benchmark for tracking model capability after release.
> 原文链接: https://github.com/ninjahawk/livenerf

---

# livenerf

[](#livenerf)

_A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch._

[](https://www.python.org/) [](https://www.anthropic.com/claude) [](https://inspect.aisi.org.uk/) [](https://code.claude.com/docs/en/headless) [](#results) [](#status)

**[📋 The plan](https://github.com/ninjahawk/livenerf/blob/main/PLAN.md)** · **[📊 Results](#results)** · **[🔬 How it works](#how-it-works)** · **[🧪 Pre-registration](#pre-registration)**


* * *

livenerf is a small, boring, append-only benchmark for one question: does a model get worse after it ships? For months there have been reports that Anthropic "nerfs" models some days or weeks after release. That could mean quantization, a smaller model behind the same name, lower effort, or routing changes. It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes. Claude Opus 5.5 came out on 2026-09-22, so this is a chance to start the clock on launch day and keep it running. Right now v0 runs on a Claude Max subscription through headless Claude Code (`claude -p`), with no API key. You can't make these models deterministic: sampling params are gone a

Notes

  • Content grounded in a same-day fetch (opencli web read / twitter thread / direct HTTP) during daily-intake-evening.
  • 中文导读 block is the entry's summary_zh (verbatim); 为什么值得关注 is the entry's one_liner; no claims beyond the fetched source text were added.
  • Source file cached at /tmp/aaif-evening/livenerf.md during the run.