跳转至

05 Training-free、自推测与长上下文

无需独立神经 drafter 或专门训练,利用跳层、自身稀疏 KV、历史检索与窗口化 MTP,应对长上下文中的 KV 读取税。

10 篇核心论文 182 页核读 更新至 2026

读完这一类,应能回答

  • 如何让 self-draft 变快而不破坏完整 target cache?
  • 长上下文中应该使用 window、retrieval、sparse KV 还是层级 proposal?
  • 历史复用带来的内存、隐私、污染与租户隔离成本是什么?

推荐阅读路线

  1. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — ACL 2024,2024。
  2. TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding — COLM 2024,2024。
  3. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — ICLR 2025,2024。
  4. LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification — arXiv preprint,2025。
  5. Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context — arXiv preprint,2026。

全部精读

年份 论文 Venue 核读页码
2024 Break the Sequential Dependency of LLM Inference Using Lookahead Decoding ICML 2024 1-20
2024 Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding ACL 2024 1-20
2024 MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding ICLR 2025 1-16
2024 REST: Retrieval-Based Speculative Decoding NAACL 2024 1-14
2024 SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications NeurIPS 2025 1-22
2024 TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding COLM 2024 1-16
2025 LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification arXiv preprint 1-19
2025 SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences arXiv preprint 1-12
2026 Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes arXiv preprint 1-18
2026 Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context arXiv preprint 1-25