05 Training-free、自推测与长上下文¶
无需独立神经 drafter 或专门训练,利用跳层、自身稀疏 KV、历史检索与窗口化 MTP,应对长上下文中的 KV 读取税。
10 篇核心论文
182 页核读
更新至 2026
读完这一类,应能回答¶
- 如何让 self-draft 变快而不破坏完整 target cache?
- 长上下文中应该使用 window、retrieval、sparse KV 还是层级 proposal?
- 历史复用带来的内存、隐私、污染与租户隔离成本是什么?
推荐阅读路线¶
- Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — ACL 2024,2024。
- TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding — COLM 2024,2024。
- MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — ICLR 2025,2024。
- LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification — arXiv preprint,2025。
- Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context — arXiv preprint,2026。