基础设施 4.0 · 优秀 2026-08-10 · 文章

Watch out for cache read costs

作者用 20-100 轮的 agent 任务对比账单:上下文随轮次增长,缓存读取在 Opus 5 与 GPT 5.6 Sol 上分别占 76% 与 82% 总成本GPT 5.6 Sol 在 272k token 处有一次长上下文重定价,单轮成本会翻倍KV cache 已压缩到 5GB 量级可下沉到 NVMe,这意味着缓存读取是利润最厚的环节

打开原文回到归档

Watch out for cache read costs

Source: <https://martinalderson.com/posts/watch-out-for-cache-read-costs/&gt;
Author: Martin Alderson
Original date: 2026-08-10
Captured: 2026-08-11 (AAIF daily-intake-evening)

中文摘要

作者用 20-100 轮的 agent 任务对比账单:上下文随轮次增长,缓存读取在 Opus 5 与 GPT 5.6 Sol 上分别占 76% 与 82% 总成本。GPT 5.6 Sol 在 272k token 处有一次“长上下文重定价”,单轮成本会翻倍。KV cache 已压缩到 5GB 量级、可下沉到 NVMe,这意味着缓存读取是利润最厚的环节。

English Summary

Martin Alderson benchmarks 20-100 turn agent runs: cache reads account for 76% (Opus 5) and 82% (GPT 5.6 Sol) of total cost as context grows. GPT 5.6 Sol triggers a long-context reprice at 272k tokens that doubles a single turn. KV cache is now compressible to ~5GB and spilling to NVMe, making cache reads the most profitable line item.

一句话

agent 长上下文的真账单在缓存读取,不在输入输出

Source Body Excerpt

Watch out for cache read costs

作者: Martin Alderson
发布时间: 2026-08-10T00:00:00.000Z
原文链接: https://martinalderson.com/posts/watch-out-for-cache-read-costs/

I know I'm guilty of just scanning OpenRouter's pricing tables and looking at input and output costs per million token. I've realised that's the wrong number to be focused on these days and cache read costs are actually far more important.

Most of your spend is likely cache reads

If you're running agentic workloads, _cache reads_ are almost certainly the biggest driver of costs. Since we've got much longer context windows, you probably need to update your mental maths to take into account what this does to pricing.

To take a hypothetical agentic session starting at 60k context length, with each tool call resulting in 500 tokens written and 5,000 tokens read, after 20 turns we get something like this:[\[1\]](#fn1)

| Model | Cache reads | Fresh input | Output | Total | | --- | --- | --- | --- | --- | | DeepSeek V4-Flash | $0.01 (18.4%) | $0.02 (72.8%) | $0.00 (8.8%) | $0.03 | | Claude Opus 5 | $1.04 (44.9%) | $1.03 (44.3%) | $0.25 (10.8%) | $2.32 | | GPT 5.6 Sol | $1.04 (48.1%) | $0.82 (38.0%) | $0.30 (13.9%) | $2.16 |

Cache reads are nearly half the bill. Now look what happens when we take the same session to 100 turns:

| Model | Cache reads | Fresh input | Output | Total | | --- | --- | --- | --- | --- | | DeepSeek V4-Flash | $0.09 (48.1%) | $0.08 (44.5%) | $0.01 (7.4%) | $0.19 | | Claude Opus 5 | $16.31 (76.4%) | $3.78 (17.7%) | $1.25 (5.9%) | $21.34 | | GPT 5.6 Sol | $29.55 (81.6%) | $4.70 (13.0%) | $1.96 (5.4%) | $36.20 |

You quickly see the issue. The main cost driver becomes cache reads - while you are only _adding_ 5.5k tokens each turn, the _existing_ context window has to be read on each turn, so the cumulative cost grows quadratically with the number of turns.[\[2\]](#fn2)

This also underscores how _reducing_ number of tool calls per run has an outsized impact on costs. If you can give the agent more specialised tools that require fewer turns, even cutting the number of turns down by 10% reduces cost per agent run by around 16%.

The case of the shrinking KV cache

While context windows have rocketed up in size, their size in memory has shrank rapidly. DeepSeek's KV cache algos (Compressed Sparse Attention and Heavily Compressed Attention), for example, allow a 1M context window at ~fp8 precision in around 5GB.

This has allowed KV cache to be offloaded to system memory and, increasingly, NVMe flash drives - explaining why the cost of NVMe has skyrocketed recently. Given the huge leaps in KV cache compression, a 1-5GB KV cache can be written to SSD and read back _extremely_ quickly, especially with RAID-style setups and PCIe 5.0 flash storage (in theory _well_ under 100ms is possible). And with both Nvidia and AMD supporting direct NVMe read and writes to the GPU, it doesn't even need to touch system RAM. As such a bank of NVMe drives can host tens of thousands of agentic sessions.

DeepSeek have made this a huge selling point of their inference API - offering cache reads at a tenth of the cost of other providers of the same model. As I'm writing this they are rumoured to be _increasing_ this price, but I'm sure this is because of huge hardware imbalances on their side, not any underlying reason. I'm sure the market will start bidding the price of cache reads down significantly.

This is (probably?) a huge profit centre

Cache reads are almost certainly outrageously profitable for the frontier labs. You're effectively paying over and over again to read a handful of GB of (V)RAM. Given most serving architectures allow you to boot this off VRAM quickly and onto system RAM (or even NVMe), you're effectively renting a few GB of system RAM at a spectacular markup.

The maths on the 100-turn run above shows us that Opus 5 spends $16.31 on cache reads. At two minutes a turn that session runs about 3.3 hours, and the context averages roughly 330k tokens over its life, so even assuming a much-larger-than-deepseek 30KB/token you're holding around 10GB. That works out at somewhere around $0.5 per GB-hour. AWS will rent you memory for well under a cent per GB-hour.

Now I'm oversimplifying here, because there _are_ definite costs to tiered KV cache storage that go beyond RAM and NVMe (such as very complex and expensive networking to make sure the KV cache is in the right place at the right time). But, if we start seeing more local and on prem LLM solutions, this cost is pretty minimal for most organisations - it only gets super complex at huge scale.

The key learning I took away from this is that cache read costs is increasingly going to be the main cost you need to look out for. There has been _huge_ innovation in making the underlying caches far, far smaller and the pricing mechanism hasn't really adjusted.

  • * *

1. YMMV significantly on this, but for many document analysis agentic tasks this matches my real world experience - the model thinks for a bit, then does a very 'short' bash command to grep through documents, _returning_ a lot of tokens. My coding sessions are similar too, with most of the time being spent grepping for existing code rather than writing it. I'm also assuming a 100% cache hit rate, which is probably reasonable for most autonomous agents. ↩︎

2. There's a second thing hiding in those two tables. At 20 turns GPT 5.6 Sol was _cheaper_ than Opus 5 ($2.16 vs $2.32). At 100 turns it's 70% more expensive ($36.20 vs $21.34). That's not a rounding artefact - OpenAI re-prices the _entire call_ at 2x input and 1.5x output once you cross 272k input tokens, so every turn after that point costs double. Anthropic explicitly doesn't do this, and says so: "a 900k-token request is billed at the same per-token rate as a 9k-token request". In the run above, that cliff alone adds 76% to GPT 5.6 Sol's bill. Which is really the whole point - the rate card told you GPT 5.6 Sol was the cheaper option, and for any session that runs long it was wrong. ↩︎

If you found this useful, I send a newsletter every month with all my posts. No spam and no ads.