基础设施 4.0 · 优秀 2026-08-28 · X

LLM 推理栈里 4 种 cache:KV / Prefix / Prompt / Semantic 存的根本不是同一类东西

@_avichawla 8月28日长文KV, Prefix, Prompt and Semantic Caching in LLMs, clearly explained从第一性原理拆分四种 cache:KV cache 存的是单次请求内每个 token 每层 attention 张量;Prefix cache 按 token ID 哈希链放服务端做跨请求复用;Prompt cache 是把 prefix cache 包成计费产品(读 0.1x写 1.25x);Semantic cache 存生成结果字符串按 embedding 余弦相似度做模糊匹配前三种是精确匹配 + 正确性中性,命中失败只花钱+拖慢...

打开原文回到归档

LLM 推理栈里 4 种 cache:KV / Prefix / Prompt / Semantic 存的根本不是同一类东西

  • ID: efc1f361
  • 原文链接: https://x.com/_avichawla/status/2093265776266637739
  • 作者/平台: @_avichawla / x
  • 发布日期: 2026-08-28
  • 归档分类: infra
  • 标签: caching、kv-cache、prompt-cache、semantic-cache、agent-harness
  • 质量评分: 4/5
  • 抓取时间: 2026-09-01T23:30+08:00

中文导读

@_avichawla 8月28日长文KV, Prefix, Prompt and Semantic Caching in LLMs, clearly explained从第一性原理拆分四种 cache:KV cache 存的是单次请求内每个 token 每层 attention 张量;Prefix cache 按 token ID 哈希链放服务端做跨请求复用;Prompt cache 是把 prefix cache 包成计费产品(读 0.1x写 1.25x);Semantic cache 存生成结果字符串按 embedding 余弦相似度做模糊匹配前三种是精确匹配 + 正确性中性,命中失败只花钱+拖慢;第四种是模糊匹配,会返回看起来对但其实错的内容且接口正常返回 200代码示例基于 360M 模型 + 单机 CPU + transformers v5,并指出 v5 的 DynamicCache() 已不再要 config 参数,dtype 取代 torch_dtype,老博客代码大概率跑不动

为什么值得关注

KV/Prefix/Prompt 三种 cache 是精确匹配,Semantic cache 是模糊匹配会返回错误答案做 Agent harness 时必须分清

关键信息

  • 文章标题:LLM 推理栈里 4 种 cache:KV / Prefix / Prompt / Semantic 存的根本不是同一类东西
  • 作者/平台:@_avichawla / x
  • 原文链接:https://x.com/_avichawla/status/2093265776266637739
  • 发布日期:2026-08-28
  • 关联标签:caching、kv-cache、prompt-cache、semantic-cache、agent-harness

English Summary

An Avi Chawla thread distinguishing four LLM caches by what they actually store and what failure looks like: KV cache (per-token attention tensors), Prefix cache (hashed token-ID chain on the server for cross-request reuse), Prompt cache (Prefix cache wrapped as a billed product with 0.1x read / 1.25x write), and Semantic cache (string embeddings matched by cosine similarity). The first three are exact-match and correctness-neutral; miss costs money and latency. Semantic cache is fuzzy-match and can return wrong-looking-right answers with HTTP 200. Code samples run on a 360M model + transformers v5, which already changed its cache API: DynamicCache() no longer takes a config arg, dtype replaces torch_dtype. Old blog code likely won't run.

Obsidian Notes

  • 来源:2026-09-01 AK-RSS Digest(89源精选)/ 每日综合摘要 / 调研 / DeepResearch 视所属主题而定
  • 内容由 opencli 拉取原始来源 + Obsidian 笔记交叉核对生成。
  • 中文导读与价值判断均锚定原文摘要与作者;未补充原文章节之外的细节。
  • 抓取时间戳:2026-09-01T23:30+08:00。