基础设施 4.0 · 优秀 2026-08-27 · 论文

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

块草拟(一次前向提出多个 token)的拒绝率混合了两种损失:块内路径信息缺失与对可观测信息建模不完美,仅看接受长度无法区分本文提出信息地板(information floor):给定条件化顺序下的最低期望拒绝率,超过地板的部分即模型差距在四个领域四个开源目标模型与一个前沿 API 模型上估计得到三个发现:Qwen3-4B 末槽全并行地板达 0.286,任何提议器单槽接受率上限 71%;一个已实现 token 可消除 86-100% 的地板(局域性,独立的互信息分析同样复原该结论);当前草拟器远高于地板:末槽模型差距占 DFlash 拒绝的 43-64%DSpark oracle 条件拒绝的 85-92%这些结果把短程条件化的价值与提议器质量区分开来

打开原文回到归档

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

  • ID: 8c39e26a
  • 原文链接: https://arxiv.org/abs/2608.27339
  • PDF: https://arxiv.org/pdf/2608.27339v1
  • 作者: Xinwei Qiang, Xiang Fang, Chang Chen, Yue Guan, Yufei Ding
  • 日期: 2026-08-27
  • 更新: 2026-08-27
  • 分类: infra
  • 来源类型: paper
  • 标签: speculative-decoding, block-drafting, information-theory, inference-optimization, acceptance-rate
  • 质量评分: 4/5
  • 抓取时间: 2026-08-31T12:36:54Z

中文导读

块草拟(一次前向提出多个 token)的拒绝率混合了两种损失:块内路径信息缺失与对可观测信息建模不完美,仅看接受长度无法区分本文提出信息地板(information floor):给定条件化顺序下的最低期望拒绝率,超过地板的部分即模型差距在四个领域四个开源目标模型与一个前沿 API 模型上估计得到三个发现:Qwen3-4B 末槽全并行地板达 0.286,任何提议器单槽接受率上限 71%;一个已实现 token 可消除 86-100% 的地板(局域性,独立的互信息分析同样复原该结论);当前草拟器远高于地板:末槽模型差距占 DFlash 拒绝的 43-64%DSpark oracle 条件拒绝的 85-92%这些结果把短程条件化的价值与提议器质量区分开来

为什么值得关注

块草拟(一次前向提出多个 token)的拒绝率混合了两种损失:块内路径信息缺失与对可观测信息建模不完美,仅看接受长度无法区分本文提出信息地板(information floor):给定条件化顺序下的最低期望拒绝率...

论文于 2026-08-27 提交至 arXiv(分类:c, s, ., L, G, ,, , c, s, ., C, L, ,, , c, s, ., I, T),arXiv 摘要页面:https://arxiv.org/abs/2608.27339。

关键信息

  • 论文标题:Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
  • 作者:Xinwei Qiang, Xiang Fang, Chang Chen, Yue Guan, Yufei Ding
  • arXiv:https://arxiv.org/abs/2608.27339
  • 发布时间:2026-08-27
  • arXiv 分类:c, s, ., L, G, ,, , c, s, ., C, L, ,, , c, s, ., I, T
  • 关联标签:speculative-decoding, block-drafting, information-theory, inference-optimization, acceptance-rate

English Abstract

Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.

English Summary

Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。