A Universal Context-Reuse Layer for Cross-Model KV Sharing
- ID: bf7e4e36
- 原文链接: https://arxiv.org/abs/2608.30963
- PDF: https://arxiv.org/pdf/2608.30963v1
- 作者: Yi Li, Dongming Jiang, Yi Zhao, Bingzhe Li
- 日期: 2026-08-31
- 抓取时间: 2026-09-02T15:31:00Z
- 分类: infra
- 来源类型: paper
- 语言: zh
- 标签: kv-cache, cross-model, inference, prefill, agent-pipeline
- 质量评分: 4/5
中文导读
现有 KV cache 复用几乎都假设生产者与消费者是同一模型,这篇把跨模型 KV 共享做成系统抽象,研究如何把 A 模型的 KV 状态翻译给规模架构attentiontokenizer模型家族都不同的 B 模型消费同族 Qwen2.5-7B Qwen2.5-1.5B 在 LongBench2 把准确率从 27.59% 提到 34.48%;跨族 Qwen2.5-1.5B Gemma-2-2B 在 4K 上下文下目标侧 prefill 成本最多降 67.05%;更异构的 Llama3.1-70B Qwen2.5-7B 拿到 44.0% vs 原生 45.7% 的准确率同时把延迟从 899ms 降到 138ms跨家族 1.7 点换 6.5 倍延迟改善的取舍对多 agent 流水线往往划算,但跨 tokenizer 的翻译边界作者自认尚未说透
一句话点评
现有 KV cache 复用几乎都假设生产者与消费者是同一模型,这篇把跨模型 KV 共享做成系统抽象,研究如何把 A 模型的 KV 状态翻译给规模架构attentiontokenizer模型家族都不同的 B 模型消费同族 Qwen2.
English Abstract / Summary
The paper studies translating KV state produced by a source model into a representation consumable by a different target model spanning scale, architecture, attention, tokenizer, and family. Within-family Qwen2.5-7B to Qwen2.5-1.5B raises LongBench2 from 27.59% to 34.48%; cross-family Qwen2.5-1.5B to Gemma-2-2B cuts target prefill cost by up to 67.05% at 4K context; Llama3.1-70B to Qwen2.5-7B keeps 44.0% accuracy versus 45.7% native while cutting latency from 899ms to 138ms. The authors frame "context movement" as a system-level abstraction and acknowledge cross-tokenizer translation boundaries remain to be characterized.
Obsidian Notes
- 由
daily-intake-evening2026-09-02 cron 从当日 Obsidian 摘要(论文流水线 / AK-RSS / ClawFeed / X 书签消化)发现并入库存量阶段。 - 中文导读与判断均锚定在条目已有摘要、源页面正文、作者、日期与分类信息;未补充源页面之外的实验细节。