On-Demand Attention: Language Models Know When to Recall
- ID: da680823
- 原文链接: https://arxiv.org/abs/2609.20734
- PDF: https://arxiv.org/pdf/2609.20734
- 作者: Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
- 日期: 2026-09-17
- 更新: 2026-09-17
- 分类: infra
- 来源类型: paper
- 标签: arxiv, paper, long-context, attention, decoding, vllm
- 质量评分: 4/5
- 抓取时间: 2026-09-21T04:30Z
中文导读
- ODA 的观察:预训练模型解码时的隐状态在全局读取之前,就已包含“读全局历史对下一步预测有多大收益”的可预测信息。
- 做法是 local-first:平时只做 local attention,轻量 recall head 按预测收益决定是否触发全局读取;只训 recall head,预训练权重不动,完整 KV cache 保留可回取。
- 作者把 GPU 条件执行做进了 vLLM,减少的全局读取转成实际解码加速(长上下文场景,优于 full attention)。
- 覆盖 Qwen 与 Gemma(含 hybrid-attention 骨干):选择性回取恢复 local attention 损失的大部分质量,同时大幅减少全局读取。
为什么值得关注
ODA 让 LLM 自学何时值得做 full attention:recall head 触发全局读,在长上下文解码上拿到实际 vLLM 加速
关键信息
- 论文标题:On-Demand Attention: Language Models Know When to Recall
- 作者:Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
- arXiv:https://arxiv.org/abs/2609.20734
- 发布时间:2026-09-17
- arXiv 分类:cs.CL
- 关联标签:arxiv, paper, long-context, attention, decoding, vllm
English Abstract
Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.
English Summary
Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.
Obsidian Notes
- 内容由
opencli arxiv paper 2609.20734 -f json拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。