Fetching the paper…
Reading the bibliography…
Speculative decoding accelerates Large Language Model (LLM) inference by using a small draft model to predict multiple tokens, and a large target model to verify these tokens in parallel.
On optimistic methods for concurrency control
H. T. Kung and John T. Robinson · 1979
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond, 2016
Ramesh Nallapati, Bowen Zhou, Cicero Nogueira dos santos, Caglar Gulcehre, and Bing Xiang · 2016
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2022
Earlier work this paper cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, L. Sifre, and John M. Jumper · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao · 2024
Earlier work this paper cites.
EAGLE: Speculative sampling requires rethinking feature uncertainty
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang · 2024
Earlier work this paper cites.
EAGLE-2: Faster inference of language models with dynamic draft trees
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang · 2024
Earlier work this paper cites.
Learning harmonized representations for speculative sampling
Lefan Zhang, Xiaodan Wang, Yanhua Huang, and Ruiwen Xu · 2024
Cited alongside, same era.
Sequoia: Scalable, robust, and hardware-aware speculative decoding
Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen · 2024
Cited alongside, same era.
Specdec++: Boosting speculative decoding via adaptive candidate lengths
Kaixuan Huang, Xudong Guo, and Mengdi Wang · 2024
Cited alongside, same era.
Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen, Vashisth Tiwari, Ruihang Lai, Jinyuan Shi, Ian En-Hsu Yen, Avner May, Tianqi Chen, and Beidi Chen · 2024
Cited alongside, same era.
Layerskip: Enabling early exit inference and self-speculative decoding
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, et al · 2024
Later among the works it cites.
Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding
Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen · 2024
Later among the works it cites.
Divide, reweight, and conquer: A logit arithmetic approach for in-context learning
Chengsong Huang, Langlin Huang, and Jiaxin Huang · 2024
Later among the works it cites.
Non-myopic generation of language models for reasoning and planning
Chang Ma, Haiteng Zhao, Junlei Zhang, Junxian He, and Lingpeng Kong · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Cited alongside, same era.
Ems-sd: Efficient multi-sample speculative decoding for accelerating large language models
Yunsheng Ni, Chuanjian Liu, Yehui Tang, Kai Han, and Yunhe Wang · 2024
Cited alongside, same era.
Turning trash into treasure: Accelerating inference of large language models with token recycling
Xianzhen Luo, Yixuan Wang, Qingfu Zhu, Zhiming Zhang, Xuanyu Zhang, Qing Yang, Dongliang Xu, and Wanxiang Che · 2024
Cited alongside, same era.
Lossless acceleration of large language model via adaptive n-gram parallel decoding
Jie Ou, Yueming Chen, and Wenhong Tian · 2024
Cited alongside, same era.
Parallel decoding via hidden transfer for lossless large language model acceleration
Pengfei Wu, Jiahao Liu, Zhuocheng Gong, Qifan Wang, Jinpeng Li, Jingang Wang, Xunliang Cai, and Dongyan Zhao · 2024
Cited alongside, same era.
Parallel speculative decoding with adaptive draft length
Tianyu Liu, Yun Li, Qitan Lv, Kai Liu, Jianchen Zhu, and Winston Hu · 2024
Cited alongside, same era.
Fast and accurate language model decoding via parallel token processing
Zhepei Wei, Wei-Lin Chen, Xinyu Zhu, and Yu Meng · 2024
Cited alongside, same era.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al · 2024
Cited alongside, same era.
Toward effective retrieval augmented generative services in 6g networks
Xi Huang, Yinxu Tang, Junling Li, Ning Zhang, and Xuemin Shen · 2024
Later among the works it cites.
Token-driven gammatune: Adaptive calibration for enchanced speculative decoding
Aayush Gautam, Susav Shrestha, and Narasimha Annapareddy · 2025
Closest in time.
Griffin: Effective token alignment for faster speculative decoding
Shijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu, Kim-Chuan Toh, and Pan Zhou · 2025
Closest in time.
Judge decoding: Faster speculative sampling requires going beyond model alignment
Gregor Bachmann, Sotiris Anagnostidis, Albert Pumarola, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Edgar Schönfeld, Ali Thabet, and Jonas Kohler · 2025
Closest in time.
Specreason: Fast and accurate inference-time compute via speculative reasoning
Rui Pan, Yinwei Dai, Zhihao Zhang, Gabriele Oliaro, Zhihao Jia, and Ravi Netravali · 2025
Closest in time.
Speculative thinking: Enhancing small-model reasoning with large model guidance at inference time
Wang Yang, Xiang Yue, Vipin Chaudhary, and Xiaotian Han · 2025
Closest in time.
Efficient reasoning for llms through speculative chain-of-thought
Jikai Wang, Juntao Li, Lijun Wu, and Min Zhang · 2025
Closest in time.
Efficient test-time scaling via self-calibration
Chengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu, and Jiaxin Huang · 2025
Closest in time.
OPT-tree: Speculative decoding with adaptive draft tree structure
Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang · 2025
Closest in time.