Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
- ID: 09495f6d
- 原文链接: https://arxiv.org/abs/2608.19920
- PDF: https://arxiv.org/pdf/2608.19920v1
- 作者: Matthias Seeger, Zeyu Zhang, Vihang Patil, Konstantinos Benidis, Sebastian Schelter
- 日期: 2026-08-20
- 分类: infra
- 来源类型: arxiv
- 代码: https://github.com/awslabs/keys_values
- 质量评分: 4/5
- 抓取时间: 2026-08-24T23:40:00+08:00
中文导读
KV 缓存选择与压缩是长上下文推理省算力的主流路线,但压缩后掉点一直靠训练时用精确注意力绕开。本文给出针对稀疏注意力的微调方法:适用于任意 KV 缓存策略,单张 A100 40GB 即可跑,让模型与驱逐策略共同适应后常能超过用精确注意力(序列并行)训练的模型。领先策略 H2O 获得带专用 SDPA kernel 的高效实现;全部方法收录于新开源库 KeysAndValues,长上下文推理与微调代码开箱即用。
Abstract (grounding)
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.
证据摘录(Obsidian 论文流水线 2026-08-24)
做长上下文推理降本的团队可以直接在这个库上对齐自己的策略再做共适应微调。