基础设施 4.0 · 优秀 2026-08-20 · 论文

Learning how to Forget: fine-tuning for long-context sparse attention

KV 缓存压缩后掉点一直靠训练时用精确注意力绕开,本文给出针对稀疏注意力的微调方法:适用任意 KV 缓存策略,单张 A100 40GB 可跑,模型与驱逐策略共同适应后常超过精确注意力训练的模型领先策略 H2O 有专用 SDPA kernel 实现,全部方法收录于 AWS 开源库 KeysAndValues

打开原文回到归档

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

中文导读

KV 缓存选择与压缩是长上下文推理省算力的主流路线,但压缩后掉点一直靠训练时用精确注意力绕开。本文给出针对稀疏注意力的微调方法:适用于任意 KV 缓存策略,单张 A100 40GB 即可跑,让模型与驱逐策略共同适应后常能超过用精确注意力(序列并行)训练的模型。领先策略 H2O 获得带专用 SDPA kernel 的高效实现;全部方法收录于新开源库 KeysAndValues,长上下文推理与微调代码开箱即用。

Abstract (grounding)

A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.

证据摘录(Obsidian 论文流水线 2026-08-24)

做长上下文推理降本的团队可以直接在这个库上对齐自己的策略再做共适应微调。