Fetching the paper…
Reading the bibliography…
Speculative decoding is an effective method for lossless acceleration of large language models during inference.
Admissibility and measurable utility functions
J. P. Quirk and R. Saposnik · 1962
Earlier work this paper cites.
Optimal transport: old and new , volume 338
C. Villani et al · 2009
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. s. Tamchyna · 2014
Earlier work this paper cites.
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
M. Stern, N. Shazeer, and J. Uszkoreit · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. Le bras, J. Gao, and Y. Choi · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Y. Leviathan, M. Kalman, and Y. Matias · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
Z. He, Z. Zhong, T. Cai, J. D. Lee, and D. He · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Inference with reference: Lossless acceleration of large language models
N. Yang, T. Ge, L. Wang, B. Jiao, D. Jiang, L. Yang, R. Majumder, and F. Wei · 2023
Later among the works it cites.
Distillspec: Improving speculative decoding via knowledge distillation, 2023
Y. Zhou, K. Lyu, A. S. Rawat, A. K. Menon, A. Rostamizadeh, S. Kumar, J.-F. Kagy, and R. Agarwal · 2023
Later among the works it cites.
Sequoia: Scalable, robust, and hardware-aware speculative decoding, 2024
Z. Chen, A. May, R. Svirschevski, Y. Huang, M. Ryabinin, Z. Jia, and B. Chen · 2024
Closest in time.
Layerskip: Enabling early exit inference and self-speculative decoding, 2024
M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, A. A. Aly, B. Chen, and C.-J. Wu · 2024
Closest in time.
Better & faster large language models via multi-token prediction, 2024
F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Online speculative decoding, 2023
X. Liu, L. Hu, P. Bailis, I. Stoica, Z. Deng, A. Cheung, and H. Zhang · 2023
Cited alongside, same era.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y. Y. Wong, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia · 2023
Cited alongside, same era.
Chatgpt-prompts
M. Rashad · 2023
Cited alongside, same era.
Sharegpt
RyokoAI · 2023
Cited alongside, same era.
Spectr: Fast speculative decoding via optimal transport
Z. Sun, A. T. Suresh, J. H. Ro, A. Beirami, H. Jain, and F. Yu · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper
Cited in the paper.
Accelerated speculative sampling based on tree monte carlo
Z. Hu and H. Huang · 2024
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty, 2024
Y. Li, F. Wei, C. Zhang, and H. Zhang · 2024
Closest in time.
Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding, 2024
H. Sun, Z. Chen, X. Yang, Y. Tian, and B. Chen · 2024
Closest in time.
Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding
H. Xia, Z. Yang, Q. Dong, P. Wang, Y. Li, T. Ge, T. Liu, W. Li, and Z. Sui · 2024
Closest in time.
Draft & verify: Lossless large language model acceleration via self-speculative decoding, 2024
J. Zhang, J. Wang, H. Li, L. Shou, K. Chen, G. Chen, and S. Mehrotra · 2024
Closest in time.