核心论文来源与核读版本¶
以下 66 项与逐篇解读一一对应。pages_read 表示本知识库核读的 PDF 页范围;SHA-256 用来确定本地核读版本。PDF 不随仓库发布。
| 年份 | 论文 | 已读版本 | 页码 | 官方入口 | PDF SHA-256 |
|---|---|---|---|---|---|
| 2018 | Blockwise Parallel Decoding for Deep Autoregressive Models | arXiv:1811.03115v1 | 1-10 | source | e34c3c9ddd1066b7dc2c11406580a586a6d483a23e6362e1414f760b655c52e1 |
| 2022 | Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation | arXiv:2203.16487v6 | 1-17 | source | 04a326ee22342c42a2bc6b1d5262053b1f42b8c321cd2463a70bf0c114c92e07 |
| 2023 | Accelerating Large Language Model Decoding with Speculative Sampling | arXiv:2302.01318v1 | 1-11 | source | ffa03c6ae46f3122570bacd7da358cae8659b6421162bbc25088622fd4889c37 |
| 2023 | Accelerating LLM Inference with Staged Speculative Decoding | arXiv:2308.04623v1 | 1-6 | source | 138febf14cb78afc03cb9e9ffac4fad48649986e60d80b4150c4e1f42ea92590 |
| 2023 | Fast Inference from Transformers via Speculative Decoding | ICML 2023 proceedings | 1-13 | source | b287adca11d8126e86dfd6facf162f74f2f6dcfefdef9130af7dde35b08e1d1d |
| 2023 | SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification | arXiv:2305.09781 / ASPLOS 2024 | 1-18 | source | 38764bfe741d39c9a309d1b715e8dc2cea6ea6c965337752b3562b1785afd9f1 |
| 2023 | SpecTr: Fast Speculative Decoding via Optimal Transport | arXiv:2310.15141v2 | 1-21 | source | 282933bead429a13c239f739376490365d724e124fc9b1a4737be27e8c80bfc4 |
| 2023 | Speculative Decoding with Big Little Decoder | arXiv:2302.07863v4 | 1-21 | source | 4a2dcdbfd818e49b20f04d0b8b52b857fe1df4504aa60fe9976919c06d018a42 |
| 2023 | The Synergy of Speculative Decoding and Batching in Serving Large Language Models | arXiv:2310.18813 | 1-9 | source | 323b330427ed0f7b6c58ea9a5fe07b82fdbf6a8051efa8a0aa9fe80223d166a7 |
| 2024 | Break the Sequential Dependency of LLM Inference Using Lookahead Decoding | ICML 2024 proceedings | 1-20 | source | b04a936f01543226ba4803a7316b4c10edea9840b0ea5aa684b926d81b4c0eb4 |
| 2024 | DistillSpec: Improving Speculative Decoding via Knowledge Distillation | arXiv:2310.08461v2 / ICLR 2024 | 1-40 | source | c9a399bf8bf5bc0658bb7212990bb589227931e85c1b80cc28e7c89924d9566e |
| 2024 | Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding | ACL 2024 proceedings | 1-20 | source | 66cfd86ab84c8a453131806c27969aad1bc63c89376cf798bee7e3650ae7fd25 |
| 2024 | EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees | EMNLP 2024 proceedings | 1-12 | source | 922fb0ff609792d6a50e43d35c45edb69a3194f4b69b1f87176c65d04ab65cad |
| 2024 | EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty | ICML 2024 proceedings | 1-14 | source | 260141e3e3942ac5797be5bd537aca30b9033d457e61e647fe119322ec9958f7 |
| 2024 | Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding | COLM 2024 paper | 1-17 | source | c531e10ad363f2498865efee1b89864bf7a5c6ad239836a80f756f0f137cb71c |
| 2024 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding | ICLR 2025 paper | 1-16 | source | a790569abe3000955becc89de7faecb613180a2e87131d3a34e2d114aa018ea5 |
| 2024 | MEDUSA: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads | ICML 2024 proceedings / arXiv:2401.10774 | 1-27 | source | 93d98f2e858c87ee04be440ee81ab8ad93700652ed759227d25fcf75cfdd5ef0 |
| 2024 | Multi-Candidate Speculative Decoding | arXiv:2401.06706 | 1-15 | source | d1875a3186dba1804766c87ce79e2c03abbe9fb3a170edf539752bf6a196474d |
| 2024 | Online Speculative Decoding | ICML 2024 proceedings | 1-16 | source | 76ab5471033a7534f24a8fbb7873d336bf61262873b19dc916c3bf2bc09ba1f4 |
| 2024 | Recurrent Drafter for Fast Speculative Decoding in Large Language Models | arXiv:2403.09919v5 | 1-14 | source | 27b266ccb64e8b0aeb25ad310e1a54fdf9709e323ecb1bc4670ffd665ee3a63a |
| 2024 | REST: Retrieval-Based Speculative Decoding | NAACL 2024 proceedings | 1-14 | source | cd810cf3d88826443cdee053329c299997fc47f4a1cdf8af895c629993230fec |
| 2024 | SEQUOIA: Scalable and Robust Speculative Decoding | arXiv:2402.12374v3 | 1-27 | source | 4e93904ae0a2a8b1813767e8a67c6709833b8affd60aff927c30619c37e4e7f8 |
| 2024 | SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices | arXiv:2406.02532 | 1-20 | source | b6af8dd38bf7e0754fefeea398342de7bdc90d5b55edd5d34e6faa0dcc59d9eb |
| 2024 | SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications | arXiv:2411.04975v3 / NeurIPS 2025 | 1-22 | source | c3cb7f16b044bbd6d2ac8849d460eba0846974b453734b2d177edae924af60f0 |
| 2024 | TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding | COLM 2024 paper | 1-16 | source | 30b4071aceff2b7925e74b979ed0f41fb227afd97862b60bb2f20700d206af3a |
| 2024 | Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding | ACL Anthology proceedings | 1-17 | source | e1514736ae0cbaa3592be40d4ec0fd3ac13d12a54866977d9c4c832bb4e49041 |
| 2025 | Block Verification Accelerates Speculative Decoding | ICLR 2025 paper | 1-30 | source | e328b76b55cebeab8cfa6a5828ac675ffc1a7137e8759c456b66259bccb7c250 |
| 2025 | Decoding Speculative Decoding | NAACL 2025 proceedings | 1-14 | source | f417b9f915fec9eb3957d832f88a8c3d943d2a74dc635f5c30295910a6b26262 |
| 2025 | EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test | NeurIPS 2025 paper | 1-20 | source | a9b7fabd038de1862791d8b73766d4a2e4caa451502f7c816215cc8e660c9277 |
| 2025 | HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding | arXiv:2505.13254 | 1-17 | source | 61d4d309f104bee98082a157cfe83aa434dcca4fefdcb23b957375031a9279bb |
| 2025 | Learning Harmonized Representations for Speculative Sampling | ICLR 2025 proceedings | 1-22 | source | 98cdd6726c24bc0ee4da1fd72acb07442a73257d19f023b9121f4e3c78e77bad |
| 2025 | LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification | arXiv:2502.17421 | 1-19 | source | f03c47105edfc62b7049130d48481689151ff07021c17184f051844e01dba637 |
| 2025 | PARD: Accelerating LLM Inference with Low-Cost Parallel Draft Model Adaptation | arXiv:2504.18583v4 | 1-18 | source | b6d8bcfd2d8a9f2ab5f96b60af72c4fb90d9f2bee86c6f0310e661d2999a938e |
| 2025 | SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences | arXiv:2505.20776 | 1-12 | source | be41fe54969f8ce94229fdc7b1d81e9e5eb46ac765ac5c7c37ccaa85114a3684 |
| 2025 | Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention | EMNLP 2025 proceedings | 1-24 | source | 521e62b5510c1d362d705e0b91f764c2f80f02824647d66cca5e84edf6769f12 |
| 2026 | Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention (SparseSpec) | MLSys 2026 proceedings | 1-15 | source | 296653f5c27d2822672e32b324b7f2e3c1ee22865091bbd4c99a44a89793b612 |
| 2026 | AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding | arXiv:2608.02989 | 1-10 | source | 771ac623b4f135ad8c191d7234881fd691c2d2c430d7443ff9c2fc6d76041c08 |
| 2026 | Adversarial Prompts for Acceptance Collapse in Speculative Decoding | arXiv:2607.21804 | 1-15 | source | e1eb3ce0e9259e9d9829404a2489de263e503db607a14e3955f72dd73151a398 |
| 2026 | AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding | arXiv:2607.25852v2 (2026-07-29) | 1-26 | source | 5eac19e3ac72136bdeab1f4d18e83c9e7a99ec34babe63514b812efebb69e324 |
| 2026 | Approximate Speculative Decoding | arXiv:2608.03447 | 1-8 | source | 76813e2e94d7be83d964df710729897e728cf7f25e9c330a3cf5aa502ff91724 |
| 2026 | CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding | arXiv:2608.00531 | 1-9 | source | 8072bf96f2aa1653a23f44041acbc1adce739fc33604e4e9cc770ac31db18189 |
| 2026 | DBLast: Dependent Block Drafting for Stochastic Speculative Decoding | arXiv:2608.05448 | 1-12 | source | ae3ef8e946c666c97578a20fa6f79863c0d420052232131b2316a496008bc5c2 |
| 2026 | DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting | arXiv:2607.07409 | 1-17 | source | 06c14c37c330f9f6f0728c2646d14258bc197f3f7ccb034707aaebca0fe9e588 |
| 2026 | DFLARE: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding | arXiv:2606.02091 | 1-12 | source | f6b301797ca5496de7bdedb15fd5ae47a04248b62074adfe283ab168a972161f |
| 2026 | DFlash: Block Diffusion for Flash Speculative Decoding | ICML 2026 proceedings | 1-13 | source | ffa514e6ce180eb1f7a39c49372f3b8170b99f8bc142d4a4daa0f087bf2ceb91 |
| 2026 | Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding | arXiv:2605.29707 | 1-11 | source | 321128dbada12c8dd7c41b497a031d1b34929d81ceb88ef95266cbba8ed24ade |
| 2026 | DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation | arXiv:2607.05147v1 (2026-07-06) | 1-33 | source | 522036b0cc16ad4678bd7c278dd0a0ab4da31170af7b97c2041067cc09a8289a |
| 2026 | From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding | arXiv:2608.02123v1 (2026-08-03) | 1-18 | source | adfade9aff11a63b1fd904f660d59af948e50c80fbc63ad47ff94752011f5f25 |
| 2026 | Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization | arXiv:2511.15898v1 / ICLR 2026 | 1-34 | source | d6144c28e5ccf1883c23e88ebc057c8526f0f1bb70d22e491f5c6c3dec0c340d |
| 2026 | JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting | arXiv:2606.18394v3 | 1-21 | source | 500750163f56a3a49939667611b63e9091a3ebdc60503bb1d626a95d0e03c142 |
| 2026 | Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware | arXiv:2607.17283v1 | 1-15 | source | 2f98f68776f01fd83197141adbaaaaa06e8406b2b81052061b807722c8d97831 |
| 2026 | MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification | arXiv:2601.15498 | 1-12 | source | a7308c5226d08bffeab845d8e59d288b977a508404c820a9d832bef1a2d6e8f9 |
| 2026 | Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding | arXiv:2605.14005v2 | 1-14 | source | fc506897f47b4f67d072b0eeda4a17af392b06dfdeb94c41202eb273d175412e |
| 2026 | Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs | arXiv:2510.20064v2 / ICLR 2026 | 1-27 | source | d8b089f553d95c46e271f792b008aa18a8abdd39d8f76e170b6a031f3427202f |
| 2026 | Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes | arXiv:2608.03839 | 1-18 | source | ffca1d97b95cd1db334023fd1b33dc4ca6336176ae651d6e8247beebc7183159 |
| 2026 | P-EAGLE: Parallel-Drafting EAGLE with Scalable Training | arXiv:2602.01469 | 1-13 | source | 35310a5280cd01e9c4d85be65aef9506728a69b17f09560ade26716a86dfbcd7 |
| 2026 | PRISM: Parametrically Refactor Inference for Speculative Decoding Draft Models | MLSys 2026 proceedings | 1-14 | source | c98885564d422a097dcc5b7f2c414a1209f3677d04f1ac072beb47f37d7534b2 |
| 2026 | Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes | arXiv:2607.26627 | 1-20 | source | e318afa118a808823c7e2613f4f3ea23bc7150ba764ad8c227f74129f9b151e3 |
| 2026 | SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts | arXiv:2608.04962 | 1-23 | source | cbad628697b64fec135c42a1acbbf38fdd20d74b148de3ed1fa8152fbf5f2e46 |
| 2026 | Speculative Decoding and the Curse of Multilinguality | arXiv:2605.30580v2 | 1-15 | source | df424a7f6e2a746f8e36eebcc6f62ec6df17b2065274079fc9fd7d920d6b4303 |
| 2026 | Speculative Decoding: Performance or Illusion? | MLSys 2026 proceedings | 1-23 | source | ca94c05c3112e46c652a17682ffe13008cb5550785f5e0fe04b743ebb67df8ca |
| 2026 | SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding | ICML 2026 proceedings / arXiv:2604.09557v2 | 1-28 | source | fedc5b8d375295148d3deb678ecd63d6e1bc144888593f58c785da5a1e5e5c12 |
| 2026 | TreeFlash: Parallel AR-Approximation for Faster Speculative Decoding | arXiv:2606.03819 | 1-13 | source | 0f41751b6ae4bf33c276ac0c646ea7c7f0620046718c25782975f7a0b78bfb39 |
| 2026 | When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding | arXiv:2606.30265v1 | 1-29 | source | bbe6f221c16e34bde65aab394ce166ccc96a702561fbeec0cf70664be89fa31e |
| 2026 | Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context | arXiv:2607.21535 | 1-25 | source | c63f01ea5839cd33710325498ec6371a5606937dab1088937d230cd73460e586 |
| 2026 | xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding | arXiv:2608.02438v1 (2026-08-03) | 1-16 | source | 3062447be736f1a4d6ee94a6ddb06c1fa98a75828035642d42009a980542b1e1 |