跳转至

01 基础、理论与综述

从 blockwise parallel decoding 到严格 speculative sampling、最优传输、多 draft 凸优化与接受率理论,建立整个方向的数学底座。

9 篇核心论文 166 页核读 更新至 2026

读完这一类,应能回答

  • 为什么一次 target forward 可以无损提交多个 token?
  • greedy-exact 与 distribution-preserving 分别需要什么条件?
  • 接受率、期望提交长度与真实 wall-clock speedup 如何区分?

推荐阅读路线

  1. Blockwise Parallel Decoding for Deep Autoregressive Models — NeurIPS 2018,2018。
  2. Fast Inference from Transformers via Speculative Decoding — ICML 2023,2023。
  3. SpecTr: Fast Speculative Decoding via Optimal Transport — NeurIPS 2023,2023。
  4. Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization — ICLR 2026 Oral,2026。

全部精读

年份 论文 Venue 核读页码
2018 Blockwise Parallel Decoding for Deep Autoregressive Models NeurIPS 2018 1-10
2022 Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation arXiv / ICLR 2023 submission 1-17
2023 Accelerating Large Language Model Decoding with Speculative Sampling arXiv technical report 1-11
2023 Fast Inference from Transformers via Speculative Decoding ICML 2023 1-13
2023 SpecTr: Fast Speculative Decoding via Optimal Transport NeurIPS 2023 1-21
2024 Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding Findings of ACL 2024 1-17
2025 Decoding Speculative Decoding NAACL 2025 1-14
2026 Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization ICLR 2026 Oral 1-34
2026 When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding arXiv preprint 1-29