跳转至

06 Serving、基准、安全与应用

把 speculative decoding 放进真实系统:dynamic batching、goodput、SLA、MoE、RL rollout、统一基准,以及输出不变但计算成本崩溃的攻击。

11 篇核心论文 190 页核读 更新至 2026

读完这一类,应能回答

  • 为什么 batch=1 的最高 speedup 不能代表生产 serving?
  • 怎样用 non-anticipating scheduler 给全 batch 分配 draft / verify 预算?
  • 如何隔离 acceptance-collapse 攻击并给出最坏额外成本上界?

推荐阅读路线

  1. The Synergy of Speculative Decoding and Batching in Serving Large Language Models — arXiv preprint,2023。
  2. Speculative Decoding: Performance or Illusion? — MLSys 2026,2026。
  3. SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding — ICML 2026,2026。
  4. Adversarial Prompts for Acceptance Collapse in Speculative Decoding — arXiv preprint,2026。

全部精读

年份 论文 Venue 核读页码
2023 The Synergy of Speculative Decoding and Batching in Serving Large Language Models arXiv preprint 1-9
2025 Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention EMNLP 2025 1-24
2026 Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention (SparseSpec) MLSys 2026 1-15
2026 AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding arXiv preprint 1-10
2026 Adversarial Prompts for Acceptance Collapse in Speculative Decoding arXiv preprint 1-15
2026 Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware arXiv preprint 1-15
2026 Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding arXiv preprint 1-14
2026 PRISM: Parametrically Refactor Inference for Speculative Decoding Draft Models MLSys 2026 1-14
2026 SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts arXiv preprint 1-23
2026 Speculative Decoding: Performance or Illusion? MLSys 2026 1-23
2026 SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding ICML 2026 1-28