基础设施 4.0 · 优秀 2026-09-17 · 论文

On-Demand Attention: Language Models Know When to Recall

提出 On-Demand Attention (ODA):给预训练模型挂一个轻量 recall head,先用 local attention 解码,只在 recall head 预测全局收益足够时再触发 global attention;推理时只训练 recall head,原模型权重不动,完整 KV cache 保留已在 Qwen / Gemma / hybrid-attention 骨干上验证:相比纯 local attention 恢复大部分质量,并显著减少 global 读次数;已移植到 vLLM GPU 条件执行,在长上下文解码场景拿到实际加速论文 28 页,2026-09-17 提交 cs.CL

打开原文回到归档

On-Demand Attention: Language Models Know When to Recall

  • ID: da680823
  • 原文链接: https://arxiv.org/abs/2609.20734
  • PDF: https://arxiv.org/pdf/2609.20734
  • 作者: Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
  • 日期: 2026-09-17
  • 更新: 2026-09-17
  • 分类: infra
  • 来源类型: paper
  • 标签: arxiv, paper, long-context, attention, decoding, vllm
  • 质量评分: 4/5
  • 抓取时间: 2026-09-21T04:30Z

中文导读

  • ODA 的观察:预训练模型解码时的隐状态在全局读取之前,就已包含“读全局历史对下一步预测有多大收益”的可预测信息。
  • 做法是 local-first:平时只做 local attention,轻量 recall head 按预测收益决定是否触发全局读取;只训 recall head,预训练权重不动,完整 KV cache 保留可回取。
  • 作者把 GPU 条件执行做进了 vLLM,减少的全局读取转成实际解码加速(长上下文场景,优于 full attention)。
  • 覆盖 Qwen 与 Gemma(含 hybrid-attention 骨干):选择性回取恢复 local attention 损失的大部分质量,同时大幅减少全局读取。

为什么值得关注

ODA 让 LLM 自学何时值得做 full attention:recall head 触发全局读,在长上下文解码上拿到实际 vLLM 加速

关键信息

  • 论文标题:On-Demand Attention: Language Models Know When to Recall
  • 作者:Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
  • arXiv:https://arxiv.org/abs/2609.20734
  • 发布时间:2026-09-17
  • arXiv 分类:cs.CL
  • 关联标签:arxiv, paper, long-context, attention, decoding, vllm

English Abstract

Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.

English Summary

Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.

Obsidian Notes

  • 内容由 opencli arxiv paper 2609.20734 -f json 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。