01 基础、理论与综述¶
从 blockwise parallel decoding 到严格 speculative sampling、最优传输、多 draft 凸优化与接受率理论,建立整个方向的数学底座。
9 篇核心论文
166 页核读
更新至 2026
读完这一类,应能回答¶
- 为什么一次 target forward 可以无损提交多个 token?
- greedy-exact 与 distribution-preserving 分别需要什么条件?
- 接受率、期望提交长度与真实 wall-clock speedup 如何区分?
推荐阅读路线¶
- Blockwise Parallel Decoding for Deep Autoregressive Models — NeurIPS 2018,2018。
- Fast Inference from Transformers via Speculative Decoding — ICML 2023,2023。
- SpecTr: Fast Speculative Decoding via Optimal Transport — NeurIPS 2023,2023。
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization — ICLR 2026 Oral,2026。
全部精读¶
| 年份 | 论文 | Venue | 核读页码 |
|---|---|---|---|
| 2018 | Blockwise Parallel Decoding for Deep Autoregressive Models | NeurIPS 2018 | 1-10 |
| 2022 | Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation | arXiv / ICLR 2023 submission | 1-17 |
| 2023 | Accelerating Large Language Model Decoding with Speculative Sampling | arXiv technical report | 1-11 |
| 2023 | Fast Inference from Transformers via Speculative Decoding | ICML 2023 | 1-13 |
| 2023 | SpecTr: Fast Speculative Decoding via Optimal Transport | NeurIPS 2023 | 1-21 |
| 2024 | Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding | Findings of ACL 2024 | 1-17 |
| 2025 | Decoding Speculative Decoding | NAACL 2025 | 1-14 |
| 2026 | Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization | ICLR 2026 Oral | 1-34 |
| 2026 | When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding | arXiv preprint | 1-29 |