The AI margin collapse is gathering pace
- ID: 61df9444
- 原文链接: https://martinalderson.com/posts/ai-margin-collapse-gathering-pace/
- 作者: Martin Alderson
- 日期: 2026-09-29
- 分类: industry
- 来源类型: article
- 标签: ai-economics, inference-cost, pricing, cache, open-weights
- 质量评分: 4/5
- 抓取时间: 2026-09-30T15:48:51Z
中文导读
推理经济系列更新:OpenAI 两个月内把 GPT-6 Luna 降价 90%(先 80% 再 50%)Sol 降 60%,Luna 一度登顶 OpenRouter 榜首;Anthropic 把 Opus 5.5 的 input/output 降 20%cache read 大降 60%,重 cache 会话实际接近半价作者的关键口径是:cache read 才是 agent 场景的真实成本重心,OpenAI 的 input/output 单价已低于 DeepSeek,但 cached input 仍远贵于对方,按会话形状比价结论完全不同(文中给出两个 agentic session 的真实账单表)判断:90% 降价难以用推理效率提升解释,旗舰大模型可能收缩为专用利基,前沿实验室要么转向推理效率与小模型,要么把大模型留作垂直整合(Anthropic 已悄悄设生物学实验室)
为什么值得关注
推理经济学最有数字的一篇:把 agent 真实成本锚在 cache read 上,并给出前沿实验室两条出路的判断。
关键信息
- 标题: The AI margin collapse is gathering pace
- 作者: Martin Alderson
- 原文链接: https://martinalderson.com/posts/ai-margin-collapse-gathering-pace/
- 发布时间: 2026-09-29
- 来源平台: blog
- 关联标签: ai-economics, inference-cost, pricing, cache, open-weights
English Summary
Latest in Alderson's inference-economics series: OpenAI cut GPT-6 Luna prices 90% in about two months (80% then a further 50%) and Sol 60%, briefly taking the OpenRouter top spot; Anthropic cut Opus 5.5 input/output by 20% and cache reads by 60%, worth nearly half-price blended on cache-heavy sessions. The pivot of the piece: cache reads, not headline input/output rates, dominate agentic cost OpenAI is now cheaper than DeepSeek on raw tokens but vastly more expensive on cached input, so session shape flips the ranking (worked-out bills for two example agentic sessions included). His read: 90% cuts are hard to explain with efficiency alone; flagships risk becoming a specialized niche, pushing labs toward inference efficiency and small models or vertical integration (cf. Anthropic's quiet biology lab).
English Source Excerpt
## Side by side
Here's how the current prices stack up (per million tokens), along with the cost of two example agentic sessions. Session A is 10M cache reads, 0.5M input and 0.2M output. Session B is more cache-heavy: 20M cache reads, 0.2M input and 0.1M output. I've ignored cache writes for simplicity.
| Model | Input | Cache read | Output | Session A | Session B |
| --- | --- | --- | --- | --- | --- |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | $0.25 | $0.27 |
| DeepSeek V4.1-Flash (off-peak) | $0.15 | $0.003 | $0.60 | $0.23 | $0.15 |
| DeepSeek V4.1-Flash (peak) | $0.30 | $0.006 | $1.20 | $0.45 | $0.30 |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | $5.00 | $5.40 |
| Opus 5 | $5.00 | $0.50 | $25.00 | $12.50 | $13.50 |
| Opus 5.5 | $4.00 | $0.20 | $20.00 | $8.00 | $6.80 |
The more cache-heavy the session, the more DeepSeek pulls ahead of Luna - and the closer Opus 5.5 gets to that 50% reduction.
## How big is the frontier, actually?
This poses a very interesting question though. Luna, Deep
Obsidian Notes
- 内容由
opencli web read抓取原文页面生成。 - 中文导读与英文摘要基于当日抓取正文与 Obsidian 摘要笔记交叉核对;英文摘录为原文逐字片段。