基础设施 4.0 · 优秀 2026-08-21 · 论文

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

开源模型架构主流正在变成混合架构:大部分层是带固定大小递归状态的 Gated DeltaNet 线性注意力层,解码时不再维护增长的 KV 缓存,但投机解码的验证-回滚需要快照完整递归状态且无法在 draft tree 分支间共享,宽树在高接受率下内存不可行TreeWY 用树状 WY 变换重写 gated delta rule:每个 draft 节点的输出用一次三角求解得到,提交时只重建被接受的那个状态,用小伪值矩阵替代逐节点快照在 Qwen3.5 35B/397B 两个规模上,相同接受长度下投机递归状态内存和 KV 缓存压力显著下降做线性注意力推理引擎的人直接拿去用

打开原文回到归档

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

中文导读

开源模型架构主流正在变成混合架构:大部分层是带固定大小递归状态的 Gated DeltaNet 线性注意力层,解码时不再维护增长的 KV 缓存,但投机解码的验证-回滚需要快照完整递归状态且无法在 draft tree 分支间共享,宽树在高接受率下内存不可行。TreeWY 用树状 WY 变换重写 gated delta rule:每个 draft 节点的输出用一次三角求解得到,提交时只重建被接受的那个状态,用小伪值矩阵替代逐节点快照。在 Qwen3.5 35B/397B 两个规模上,相同接受长度下投机递归状态内存和 KV 缓存压力显著下降。做线性注意力推理引擎的人直接拿去用。

为什么值得关注

开源模型架构主流正在变成混合架构:大部分层是带固定大小递归状态的 Gated DeltaNet 线性注意力层,解码时不再维护增长的 KV 缓存,但投机解码的验证-回滚需要快照完整递归状态且无法在 draft tree 分支间共享,宽树在高接受率下内存不可行。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:TreeWY: Speculative Verification for Gated DeltaNet Hybrids
  • 作者:Sneha Murthy Ghantasala
  • arXiv:https://arxiv.org/abs/2608.20961
  • 发布时间:2026-08-21
  • arXiv 分类:cs.AI, cs.CL, cs.DC, cs.LG, cs.PF

English Abstract

Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. This makes ordinary decoding memory-efficient, but hurts speculative decoding. To verify a batch of draft tokens and then roll back the rejected ones, today's systems snapshot the full recurrent state at every draft position for GDN layers, and those snapshots cannot be shared across branches of a draft tree, so a wide, high-acceptance tree becomes memory-infeasible. We remove the snapshots. Using a tree-structured WY transform of the gated delta rule, we compute every draft node's output with a single triangular solve and reconstruct only the one accepted state on commit, storing a small pseudo-value matrix instead of per-node states; the derivation depends only on the gated delta rule, not on any other architectural detail. In serving benchmarks on two scales of one hybrid model family (Qwen3.5 35B and 397B) this cuts speculative recurrent-state memory and KV-cache pressure at identical acceptance length, turning the freed HBM into higher throughput and much lower time-to-first-token (TTFT) wherever memory binds, and costing a few percent where it does not. For tree width the same memory buys affordability: a wider, higher-acceptance draft becomes possible, though not yet a throughput win.

English Summary

Rewrites the gated delta rule with a tree-structured WY transform so speculative-decoding verification for Gated DeltaNet hybrids needs no per-node recurrent-state snapshots: each draft node's output comes from one triangular solve, and only the accepted state is rebuilt at commit. On Qwen3.5 35B/397B, speculative recurrent-state memory and KV-cache pressure drop significantly at equal acceptance length.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。