Skip to content

Latest commit

 

History

History
72 lines (70 loc) · 20.4 KB

File metadata and controls

72 lines (70 loc) · 20.4 KB

核心论文来源与核读版本

以下 66 项与逐篇解读一一对应。pages_read 表示本知识库核读的 PDF 页范围;SHA-256 用来确定本地核读版本。PDF 不随仓库发布。

年份 论文 已读版本 页码 官方入口 PDF SHA-256
2018 Blockwise Parallel Decoding for Deep Autoregressive Models arXiv:1811.03115v1 1-10 source e34c3c9ddd1066b7dc2c11406580a586a6d483a23e6362e1414f760b655c52e1
2022 Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation arXiv:2203.16487v6 1-17 source 04a326ee22342c42a2bc6b1d5262053b1f42b8c321cd2463a70bf0c114c92e07
2023 Accelerating Large Language Model Decoding with Speculative Sampling arXiv:2302.01318v1 1-11 source ffa03c6ae46f3122570bacd7da358cae8659b6421162bbc25088622fd4889c37
2023 Accelerating LLM Inference with Staged Speculative Decoding arXiv:2308.04623v1 1-6 source 138febf14cb78afc03cb9e9ffac4fad48649986e60d80b4150c4e1f42ea92590
2023 Fast Inference from Transformers via Speculative Decoding ICML 2023 proceedings 1-13 source b287adca11d8126e86dfd6facf162f74f2f6dcfefdef9130af7dde35b08e1d1d
2023 SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification arXiv:2305.09781 / ASPLOS 2024 1-18 source 38764bfe741d39c9a309d1b715e8dc2cea6ea6c965337752b3562b1785afd9f1
2023 SpecTr: Fast Speculative Decoding via Optimal Transport arXiv:2310.15141v2 1-21 source 282933bead429a13c239f739376490365d724e124fc9b1a4737be27e8c80bfc4
2023 Speculative Decoding with Big Little Decoder arXiv:2302.07863v4 1-21 source 4a2dcdbfd818e49b20f04d0b8b52b857fe1df4504aa60fe9976919c06d018a42
2023 The Synergy of Speculative Decoding and Batching in Serving Large Language Models arXiv:2310.18813 1-9 source 323b330427ed0f7b6c58ea9a5fe07b82fdbf6a8051efa8a0aa9fe80223d166a7
2024 Break the Sequential Dependency of LLM Inference Using Lookahead Decoding ICML 2024 proceedings 1-20 source b04a936f01543226ba4803a7316b4c10edea9840b0ea5aa684b926d81b4c0eb4
2024 DistillSpec: Improving Speculative Decoding via Knowledge Distillation arXiv:2310.08461v2 / ICLR 2024 1-40 source c9a399bf8bf5bc0658bb7212990bb589227931e85c1b80cc28e7c89924d9566e
2024 Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding ACL 2024 proceedings 1-20 source 66cfd86ab84c8a453131806c27969aad1bc63c89376cf798bee7e3650ae7fd25
2024 EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees EMNLP 2024 proceedings 1-12 source 922fb0ff609792d6a50e43d35c45edb69a3194f4b69b1f87176c65d04ab65cad
2024 EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty ICML 2024 proceedings 1-14 source 260141e3e3942ac5797be5bd537aca30b9033d457e61e647fe119322ec9958f7
2024 Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding COLM 2024 paper 1-17 source c531e10ad363f2498865efee1b89864bf7a5c6ad239836a80f756f0f137cb71c
2024 MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding ICLR 2025 paper 1-16 source a790569abe3000955becc89de7faecb613180a2e87131d3a34e2d114aa018ea5
2024 MEDUSA: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads ICML 2024 proceedings / arXiv:2401.10774 1-27 source 93d98f2e858c87ee04be440ee81ab8ad93700652ed759227d25fcf75cfdd5ef0
2024 Multi-Candidate Speculative Decoding arXiv:2401.06706 1-15 source d1875a3186dba1804766c87ce79e2c03abbe9fb3a170edf539752bf6a196474d
2024 Online Speculative Decoding ICML 2024 proceedings 1-16 source 76ab5471033a7534f24a8fbb7873d336bf61262873b19dc916c3bf2bc09ba1f4
2024 Recurrent Drafter for Fast Speculative Decoding in Large Language Models arXiv:2403.09919v5 1-14 source 27b266ccb64e8b0aeb25ad310e1a54fdf9709e323ecb1bc4670ffd665ee3a63a
2024 REST: Retrieval-Based Speculative Decoding NAACL 2024 proceedings 1-14 source cd810cf3d88826443cdee053329c299997fc47f4a1cdf8af895c629993230fec
2024 SEQUOIA: Scalable and Robust Speculative Decoding arXiv:2402.12374v3 1-27 source 4e93904ae0a2a8b1813767e8a67c6709833b8affd60aff927c30619c37e4e7f8
2024 SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices arXiv:2406.02532 1-20 source b6af8dd38bf7e0754fefeea398342de7bdc90d5b55edd5d34e6faa0dcc59d9eb
2024 SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications arXiv:2411.04975v3 / NeurIPS 2025 1-22 source c3cb7f16b044bbd6d2ac8849d460eba0846974b453734b2d177edae924af60f0
2024 TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding COLM 2024 paper 1-16 source 30b4071aceff2b7925e74b979ed0f41fb227afd97862b60bb2f20700d206af3a
2024 Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding ACL Anthology proceedings 1-17 source e1514736ae0cbaa3592be40d4ec0fd3ac13d12a54866977d9c4c832bb4e49041
2025 Block Verification Accelerates Speculative Decoding ICLR 2025 paper 1-30 source e328b76b55cebeab8cfa6a5828ac675ffc1a7137e8759c456b66259bccb7c250
2025 Decoding Speculative Decoding NAACL 2025 proceedings 1-14 source f417b9f915fec9eb3957d832f88a8c3d943d2a74dc635f5c30295910a6b26262
2025 EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test NeurIPS 2025 paper 1-20 source a9b7fabd038de1862791d8b73766d4a2e4caa451502f7c816215cc8e660c9277
2025 HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding arXiv:2505.13254 1-17 source 61d4d309f104bee98082a157cfe83aa434dcca4fefdcb23b957375031a9279bb
2025 Learning Harmonized Representations for Speculative Sampling ICLR 2025 proceedings 1-22 source 98cdd6726c24bc0ee4da1fd72acb07442a73257d19f023b9121f4e3c78e77bad
2025 LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification arXiv:2502.17421 1-19 source f03c47105edfc62b7049130d48481689151ff07021c17184f051844e01dba637
2025 PARD: Accelerating LLM Inference with Low-Cost Parallel Draft Model Adaptation arXiv:2504.18583v4 1-18 source b6d8bcfd2d8a9f2ab5f96b60af72c4fb90d9f2bee86c6f0310e661d2999a938e
2025 SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences arXiv:2505.20776 1-12 source be41fe54969f8ce94229fdc7b1d81e9e5eb46ac765ac5c7c37ccaa85114a3684
2025 Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention EMNLP 2025 proceedings 1-24 source 521e62b5510c1d362d705e0b91f764c2f80f02824647d66cca5e84edf6769f12
2026 Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention (SparseSpec) MLSys 2026 proceedings 1-15 source 296653f5c27d2822672e32b324b7f2e3c1ee22865091bbd4c99a44a89793b612
2026 AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding arXiv:2608.02989 1-10 source 771ac623b4f135ad8c191d7234881fd691c2d2c430d7443ff9c2fc6d76041c08
2026 Adversarial Prompts for Acceptance Collapse in Speculative Decoding arXiv:2607.21804 1-15 source e1eb3ce0e9259e9d9829404a2489de263e503db607a14e3955f72dd73151a398
2026 AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding arXiv:2607.25852v2 (2026-07-29) 1-26 source 5eac19e3ac72136bdeab1f4d18e83c9e7a99ec34babe63514b812efebb69e324
2026 Approximate Speculative Decoding arXiv:2608.03447 1-8 source 76813e2e94d7be83d964df710729897e728cf7f25e9c330a3cf5aa502ff91724
2026 CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding arXiv:2608.00531 1-9 source 8072bf96f2aa1653a23f44041acbc1adce739fc33604e4e9cc770ac31db18189
2026 DBLast: Dependent Block Drafting for Stochastic Speculative Decoding arXiv:2608.05448 1-12 source ae3ef8e946c666c97578a20fa6f79863c0d420052232131b2316a496008bc5c2
2026 DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting arXiv:2607.07409 1-17 source 06c14c37c330f9f6f0728c2646d14258bc197f3f7ccb034707aaebca0fe9e588
2026 DFLARE: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding arXiv:2606.02091 1-12 source f6b301797ca5496de7bdedb15fd5ae47a04248b62074adfe283ab168a972161f
2026 DFlash: Block Diffusion for Flash Speculative Decoding ICML 2026 proceedings 1-13 source ffa514e6ce180eb1f7a39c49372f3b8170b99f8bc142d4a4daa0f087bf2ceb91
2026 Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding arXiv:2605.29707 1-11 source 321128dbada12c8dd7c41b497a031d1b34929d81ceb88ef95266cbba8ed24ade
2026 DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation arXiv:2607.05147v1 (2026-07-06) 1-33 source 522036b0cc16ad4678bd7c278dd0a0ab4da31170af7b97c2041067cc09a8289a
2026 From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding arXiv:2608.02123v1 (2026-08-03) 1-18 source adfade9aff11a63b1fd904f660d59af948e50c80fbc63ad47ff94752011f5f25
2026 Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization arXiv:2511.15898v1 / ICLR 2026 1-34 source d6144c28e5ccf1883c23e88ebc057c8526f0f1bb70d22e491f5c6c3dec0c340d
2026 JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting arXiv:2606.18394v3 1-21 source 500750163f56a3a49939667611b63e9091a3ebdc60503bb1d626a95d0e03c142
2026 Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware arXiv:2607.17283v1 1-15 source 2f98f68776f01fd83197141adbaaaaa06e8406b2b81052061b807722c8d97831
2026 MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification arXiv:2601.15498 1-12 source a7308c5226d08bffeab845d8e59d288b977a508404c820a9d832bef1a2d6e8f9
2026 Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding arXiv:2605.14005v2 1-14 source fc506897f47b4f67d072b0eeda4a17af392b06dfdeb94c41202eb273d175412e
2026 Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs arXiv:2510.20064v2 / ICLR 2026 1-27 source d8b089f553d95c46e271f792b008aa18a8abdd39d8f76e170b6a031f3427202f
2026 Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes arXiv:2608.03839 1-18 source ffca1d97b95cd1db334023fd1b33dc4ca6336176ae651d6e8247beebc7183159
2026 P-EAGLE: Parallel-Drafting EAGLE with Scalable Training arXiv:2602.01469 1-13 source 35310a5280cd01e9c4d85be65aef9506728a69b17f09560ade26716a86dfbcd7
2026 PRISM: Parametrically Refactor Inference for Speculative Decoding Draft Models MLSys 2026 proceedings 1-14 source c98885564d422a097dcc5b7f2c414a1209f3677d04f1ac072beb47f37d7534b2
2026 Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes arXiv:2607.26627 1-20 source e318afa118a808823c7e2613f4f3ea23bc7150ba764ad8c227f74129f9b151e3
2026 SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts arXiv:2608.04962 1-23 source cbad628697b64fec135c42a1acbbf38fdd20d74b148de3ed1fa8152fbf5f2e46
2026 Speculative Decoding and the Curse of Multilinguality arXiv:2605.30580v2 1-15 source df424a7f6e2a746f8e36eebcc6f62ec6df17b2065274079fc9fd7d920d6b4303
2026 Speculative Decoding: Performance or Illusion? MLSys 2026 proceedings 1-23 source ca94c05c3112e46c652a17682ffe13008cb5550785f5e0fe04b743ebb67df8ca
2026 SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding ICML 2026 proceedings / arXiv:2604.09557v2 1-28 source fedc5b8d375295148d3deb678ecd63d6e1bc144888593f58c785da5a1e5e5c12
2026 TreeFlash: Parallel AR-Approximation for Faster Speculative Decoding arXiv:2606.03819 1-13 source 0f41751b6ae4bf33c276ac0c646ea7c7f0620046718c25782975f7a0b78bfb39
2026 When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding arXiv:2606.30265v1 1-29 source bbe6f221c16e34bde65aab394ce166ccc96a702561fbeec0cf70664be89fa31e
2026 Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context arXiv:2607.21535 1-25 source c63f01ea5839cd33710325498ec6371a5606937dab1088937d230cd73460e586
2026 xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding arXiv:2608.02438v1 (2026-08-03) 1-16 source 3062447be736f1a4d6ee94a6ddb06c1fa98a75828035642d42009a980542b1e1