产品与商业 4.0 · 优秀 2026-09-29 · 文章

The AI margin collapse is gathering pace

推理经济系列更新:OpenAI 两个月内把 GPT-6 Luna 降价 90%(先 80% 再 50%)Sol 降 60%,Luna 一度登顶 OpenRouter 榜首;Anthropic 把 Opus 5.5 的 input/output 降 20%cache read 大降 60%,重 cache 会话实际接近半价作者的关键口径是:cache read 才是 agent 场景的真实成本重心,OpenAI 的 input/output 单价已低于 DeepSeek,但 cached input 仍远贵于对方,按会话形状比价结论完全不同(文中给出两个 agentic session 的真实账单表)判断:90% 降价难以用推理效率提升解释,旗舰大模型可能收缩为专用利基,前沿实验室要么转向推理效率与小模型...

打开原文回到归档

The AI margin collapse is gathering pace

中文导读

推理经济系列更新:OpenAI 两个月内把 GPT-6 Luna 降价 90%(先 80% 再 50%)Sol 降 60%,Luna 一度登顶 OpenRouter 榜首;Anthropic 把 Opus 5.5 的 input/output 降 20%cache read 大降 60%,重 cache 会话实际接近半价作者的关键口径是:cache read 才是 agent 场景的真实成本重心,OpenAI 的 input/output 单价已低于 DeepSeek,但 cached input 仍远贵于对方,按会话形状比价结论完全不同(文中给出两个 agentic session 的真实账单表)判断:90% 降价难以用推理效率提升解释,旗舰大模型可能收缩为专用利基,前沿实验室要么转向推理效率与小模型,要么把大模型留作垂直整合(Anthropic 已悄悄设生物学实验室)

为什么值得关注

推理经济学最有数字的一篇:把 agent 真实成本锚在 cache read 上,并给出前沿实验室两条出路的判断。

关键信息

English Summary

Latest in Alderson's inference-economics series: OpenAI cut GPT-6 Luna prices 90% in about two months (80% then a further 50%) and Sol 60%, briefly taking the OpenRouter top spot; Anthropic cut Opus 5.5 input/output by 20% and cache reads by 60%, worth nearly half-price blended on cache-heavy sessions. The pivot of the piece: cache reads, not headline input/output rates, dominate agentic cost OpenAI is now cheaper than DeepSeek on raw tokens but vastly more expensive on cached input, so session shape flips the ranking (worked-out bills for two example agentic sessions included). His read: 90% cuts are hard to explain with efficiency alone; flagships risk becoming a specialized niche, pushing labs toward inference efficiency and small models or vertical integration (cf. Anthropic's quiet biology lab).

English Source Excerpt

## Side by side
Here's how the current prices stack up (per million tokens), along with the cost of two example agentic sessions. Session A is 10M cache reads, 0.5M input and 0.2M output. Session B is more cache-heavy: 20M cache reads, 0.2M input and 0.1M output. I've ignored cache writes for simplicity.
| Model | Input | Cache read | Output | Session A | Session B |
| --- | --- | --- | --- | --- | --- |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | $0.25 | $0.27 |
| DeepSeek V4.1-Flash (off-peak) | $0.15 | $0.003 | $0.60 | $0.23 | $0.15 |
| DeepSeek V4.1-Flash (peak) | $0.30 | $0.006 | $1.20 | $0.45 | $0.30 |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | $5.00 | $5.40 |
| Opus 5 | $5.00 | $0.50 | $25.00 | $12.50 | $13.50 |
| Opus 5.5 | $4.00 | $0.20 | $20.00 | $8.00 | $6.80 |
The more cache-heavy the session, the more DeepSeek pulls ahead of Luna - and the closer Opus 5.5 gets to that 50% reduction.
## How big is the frontier, actually?
This poses a very interesting question though. Luna, Deep

Obsidian Notes

  • 内容由 opencli web read 抓取原文页面生成。
  • 中文导读与英文摘要基于当日抓取正文与 Obsidian 摘要笔记交叉核对;英文摘录为原文逐字片段。