Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- ID: 6496bc27
- 原文链接: https://arxiv.org/abs/2510.03215
- PDF: https://arxiv.org/pdf/2510.03215v2
- 作者: Tianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan, Guohao Dai, Wanli Ouyang, Yu Wang
- 日期: 2025-10-03
- 更新: 2026-03-02
- 分类: models
- 来源类型: paper
- 标签: kv-cache, multi-llm, model-communication, semantic-transfer, efficiency
- 质量评分: 4/5
- 抓取时间: 2026-09-19T04:21:30Z
- arXiv 备注: Published in ICLR'26
中文导读
多 LLM 协作的现有设计都靠文本通信:内部表示被压成 token 序列,既丢语义又串行慢C2C 提出 KV-Cache 直连:用神经网络把源模型的 KV-cache 投影并融合进目标模型,可学习门控挑选受益层,实现模型间直接语义传输实验:比单模型平均准确率高 6.4-14.2%,比文本通信高约 3.1-5.4%,延迟平均提速 2.5 倍Oracle 实验表明不增大缓存的前提下丰富 KV-Cache 语义本身就能提升回答质量ICLR'26 已发表,代码开源
为什么值得关注
LLM 之间别再传文本了:KV-Cache 直连,准确率更高还快 2.5 倍
关键信息
- 论文标题:Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- 作者:Tianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan, Guohao Dai, Wanli Ouyang, Yu Wang
- arXiv:https://arxiv.org/abs/2510.03215
- 发布时间:2025-10-03
- arXiv 分类:cs.CL, cs.LG
- 关联标签:kv-cache, multi-llm, model-communication, semantic-transfer, efficiency
English Abstract
Multi-LLM systems harness the complementary strengths of diverse Large Language Models, achieving performance and efficiency gains that are not attainable by a single model. In existing designs, LLMs communicate through text, forcing internal representations to be transformed into output token sequences. This process both loses rich semantic information and incurs token-by-token generation latency. Motivated by these limitations, we ask: Can LLMs communicate beyond text? Oracle experiments show that enriching the KV-Cache semantics can improve response quality without increasing cache size, supporting KV-Cache as an effective medium for inter-model communication. Thus, we propose Cache-to-Cache (C2C), a new paradigm for direct semantic communication between LLMs. C2C uses a neural network to project and fuse the source model's KV-cache with that of the target model to enable direct semantic transfer. A learnable gating mechanism selects the target layers that benefit from cache communication. Compared with text communication, C2C utilizes the deep, specialized semantics from both models, while avoiding explicit intermediate text generation. Experiments show that C2C achieves 6.4-14.2% higher average accuracy than individual models. It further outperforms the text communication paradigm by approximately 3.1-5.4%, while delivering an average 2.5x speedup in latency. Our code is available at https://github.com/thu-nics/C2C.
English Summary
Multi-LLM systems normally communicate through text, which loses rich semantic information and incurs token-by-token generation latency. Cache-to-Cache (C2C) proposes direct semantic communication: a neural network projects and fuses the source model's KV-cache into the target model's, with a learnable gating mechanism selecting which target layers benefit. Experiments show C2C achieves 6.4-14.2% higher average accuracy than individual models, outperforms text communication by roughly 3.1-5.4%, and delivers an average 2.5x latency speedup. Published at ICLR'26; code at github.com/thu-nics/C2C.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。