Fetching the paper…
Reading the bibliography…
Speculative decoding is a prominent technique to speed up the inference of a large target language model based on predictions of an auxiliary draft model.
Speculative computation, parallelism, and functional programming
Burton, F. W · 1985
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., Xiong, C., and Socher, R · 2017
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al · 2018
Earlier work this paper cites.
Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge
Dušek, O., Novikova, J., and Rieser, V · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Qi, W., Yan, Y., Gong, Y., Liu, D., Duan, N., Chen, J., Zhang, R., and Zhou, M · 2020
Earlier work this paper cites.
DialogSum: A real-life scenario dialogue summarization dataset
Chen, Y., Liu, Y., Chen, L., and Zhang, Y · 2021
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Breaking the sequential dependency of llm inference using lookahead decoding, November 2023
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2023
Later among the works it cites.
Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5
Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T · 2023
Later among the works it cites.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2023
Cited alongside, same era.
Medusa: Simple framework for accelerating llm generation with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., and Dao, T · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z
Cited in the paper.
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Pal, K., Sun, J., Yuan, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Later among the works it cites.
Accelerating llm inference with staged speculative decoding
Spector, B. and Re, C · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2023
Later among the works it cites.
Fastertransformer, 2024
Nvidia · 2024
Closest in time.