跳转至

02 独立 drafter、对齐与在线选择

研究独立小模型如何与 target 对齐、蒸馏、在线更新和动态选择,并揭示跨语言、跨领域与模型族迁移的边界。

7 篇核心论文 147 页核读 更新至 2026

读完这一类,应能回答

  • 什么训练目标真正优化接受率,而不只是 next-token accuracy?
  • 什么时候在线适应的收益能覆盖更新成本?
  • 一个 drafter 能否跨 checkpoint、语言和 tokenizer 迁移?

推荐阅读路线

  1. Speculative Decoding with Big Little Decoder — NeurIPS 2023,2023。
  2. DistillSpec: Improving Speculative Decoding via Knowledge Distillation — ICLR 2024,2024。
  3. Learning Harmonized Representations for Speculative Sampling — ICLR 2025,2025。
  4. Speculative Decoding and the Curse of Multilinguality — arXiv preprint,2026。

全部精读

年份 论文 Venue 核读页码
2023 Accelerating LLM Inference with Staged Speculative Decoding ICML 2023 workshop / arXiv 1-6
2023 Speculative Decoding with Big Little Decoder NeurIPS 2023 1-21
2024 DistillSpec: Improving Speculative Decoding via Knowledge Distillation ICLR 2024 1-40
2024 Online Speculative Decoding ICML 2024 1-16
2025 Learning Harmonized Representations for Speculative Sampling ICLR 2025 1-22
2026 Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs ICLR 2026 1-27
2026 Speculative Decoding and the Curse of Multilinguality arXiv preprint 1-15