02 独立 drafter、对齐与在线选择¶
研究独立小模型如何与 target 对齐、蒸馏、在线更新和动态选择,并揭示跨语言、跨领域与模型族迁移的边界。
7 篇核心论文
147 页核读
更新至 2026
读完这一类,应能回答¶
- 什么训练目标真正优化接受率,而不只是 next-token accuracy?
- 什么时候在线适应的收益能覆盖更新成本?
- 一个 drafter 能否跨 checkpoint、语言和 tokenizer 迁移?
推荐阅读路线¶
- Speculative Decoding with Big Little Decoder — NeurIPS 2023,2023。
- DistillSpec: Improving Speculative Decoding via Knowledge Distillation — ICLR 2024,2024。
- Learning Harmonized Representations for Speculative Sampling — ICLR 2025,2025。
- Speculative Decoding and the Curse of Multilinguality — arXiv preprint,2026。
全部精读¶
| 年份 | 论文 | Venue | 核读页码 |
|---|---|---|---|
| 2023 | Accelerating LLM Inference with Staged Speculative Decoding | ICML 2023 workshop / arXiv | 1-6 |
| 2023 | Speculative Decoding with Big Little Decoder | NeurIPS 2023 | 1-21 |
| 2024 | DistillSpec: Improving Speculative Decoding via Knowledge Distillation | ICLR 2024 | 1-40 |
| 2024 | Online Speculative Decoding | ICML 2024 | 1-16 |
| 2025 | Learning Harmonized Representations for Speculative Sampling | ICLR 2025 | 1-22 |
| 2026 | Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs | ICLR 2026 | 1-27 |
| 2026 | Speculative Decoding and the Curse of Multilinguality | arXiv preprint | 1-15 |