模型与实验室 4.0 · 优秀 2025-05-30 · 论文

R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

Zefan Cai 等 2025-05-30 提交,2026-01-22 更新到 v4推理模型 CoT 输出极长,KV cache 跟着爆炸;已有压缩方法在推理模型上常把推理链一起压坏R-KV 专攻推理模型里的冗余 token识别对后续推理步骤贡献小可丢弃的 token实验结果:只用 10% 的 KV cache 几乎保满性能(基线用 10% 只剩 60%),16% 时反而比满 cache 还略好(105%),配套收益是 90% 内存节省6.6 吞吐和 SGD-KV 合用,是 2026 H2 端侧长链推理的实用组合

打开原文回到归档

R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

Source: https://arxiv.org/abs/2505.24133
Authors: Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Anima Anandkumar, Abedelkadir Asi, Junjie Hu
Published: 2025-05-30
Categories: cs.CL, cs.AI
PDF: https://arxiv.org/pdf/2505.24133v1

Abstract

Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.

中文概要

Zefan Cai 等 2025-05-30 提交,2026-01-22 更新到 v4。推理模型 CoT 输出极长,KV cache 跟着爆炸;已有压缩方法在推理模型上常把推理链一起压坏。R-KV 专攻推理模型里的『冗余 token』——识别对后续推理步骤贡献小、可丢弃的 token。实验结果:只用 10% 的 KV cache 几乎保满性能(基线用 10% 只剩 60%),16% 时反而比满 cache 还略好(105%),配套收益是 90% 内存节省、6.6× 吞吐。和 SGD-KV 合用,是 2026 H2 端侧长链推理的实用组合。

一句话

R-KV 专攻推理模型冗余 token:10% KV cache 保满性能、16% 略超满 cache,端侧长链推理组合的另一块