Fetching the paper…
Reading the bibliography…
Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., and Radev, D · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Order-agnostic cross entropy for non-autoregressive machine translation
Du, C., Tu, Z., and Jiang, J · 2021
Earlier work this paper cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V. Q., Tay, Y., and Metzler, D · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
Finance-alpaca dataset, 2023
Bharti, G · 2023
Earlier work this paper cites.
Medusa: Simple framework for accelerating llm generation with multiple decoding heads, 2023
Cai, T., Li, Y., Geng, Z., Peng, H., and Dao, T · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. M · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Leveraging large language models for pre-trained recommender systems
Chu, Z., Hao, H., Ouyang, X., Wang, S., Wang, Y., Shen, Y., Gu, J., Cui, Q., Li, L., Xue, S., et al · 2023
Cited alongside, same era.
Breaking the sequential dependency of llm inference using lookahead decoding, 2023
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2023
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Later among the works it cites.
Liu, X., Hu, L., Bailis, P., Stoica, I., Deng, Z., Cheung, A., and Zhang, H · 2023
Later among the works it cites.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 2023
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
Spectr: Fast speculative decoding via optimal transport
Sun, Z., Suresh, A. T., Ro, J. H., Beirami, A., Jain, H., Yu, F., Riley, M., and Kumar, S · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sharegpt dataset, 2023
GPT3.5 and 4, G · 2023
Cited alongside, same era.
Rest: Retrieval-based speculative decoding, 2023
He, Z., Zhong, Z., Cai, T., Lee, J. D., and He, D · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Is chatgpt a good translator? yes with gpt-4 as the engine
Jiao, W., Wang, W., Huang, J.-t., Wang, X., and Tu, Z · 2023
Cited alongside, same era.
Speculative decoding with big little decoder
Kim, S., Mangalam, K., Moon, S., Malik, J., Mahoney, M. W., Gholami, A., and Keutzer, K · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Later among the works it cites.
Inference with reference: Lossless acceleration of large language models
Yang, N., Ge, T., Wang, L., Jiao, B., Jiang, D., Yang, L., Majumder, R., and Wei, F · 2023
Later among the works it cites.
Speculative contrastive decoding
Yuan, H., Lu, K., Huang, F., Yuan, Z., and Zhou, C · 2023
Later among the works it cites.
Draft and verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2024
Closest in time.