Fetching the paper…
Reading the bibliography…
Inference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al · 2016
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
The state of sparsity in deep neural networks.(2019)
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Q8bert: Quantized 8bit bert
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Earlier work this paper cites.
Instantaneous grammatical error correction with shallow aggressive decoding
Sun, X., Ge, T., Wei, F., and Wang, H · 2021
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Breaking the sequential dependency of llm inference using lookahead decoding, November 2023
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2023
Cited alongside, same era.
Speed: Speculative pipelined execution for efficient decoding
Hooper, C., Kim, S., Mohammadzadeh, H., Genc, H., Keutzer, K., Gholami, A., and Shao, S · 2023
Cited alongside, same era.
Assisted generation: a new direction toward low-latency text generation, 2023
Joao Gante · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Closest in time.
Sequoia: Scalable, robust, and hardware-aware speculative decoding
Chen, Z., May, A., Svirschevski, R., Huang, Y., Ryabinin, M., Jia, Z., and Chen, B · 2024
Closest in time.
Glide with a cape: A low-hassle method to accelerate speculative decoding
Du, C., Jiang, J., Yuanchen, X., Wu, J., Yu, S., Li, Y., Li, S., Xu, K., Nie, L., Tu, Z., et al · 2024
Closest in time.
Layer skip: Enabling early exit inference and self-speculative decoding
Elhoushi, M., Shrivastava, A., Liskovich, D., Hosmer, B., Wasti, B., Lai, L., Mahmoud, A., Acun, B., Agarwal, S., Roman, A., et al · 2024
Closest in time.
Break the sequential dependency of llm inference using lookahead decoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pass: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E · 2023
Cited alongside, same era.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R · 2023
Cited alongside, same era.
Accelerating transformer inference for translation via parallel decoding
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E · 2023
Cited alongside, same era.
Prompt lookup decoding, November 2023
Saxena, A · 2023
Cited alongside, same era.
Accelerating llm inference with staged speculative decoding
Spector, B. and Re, C · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models (2023)
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Closest in time.
Speculative decoding with big little decoder
Kim, S., Mangalam, K., Moon, S., Malik, J., Mahoney, M. W., Gholami, A., and Keutzer, K · 2024
Closest in time.
Cllms: Consistency large language models
Kou, S., Hu, L., He, Z., Deng, Z., and Zhang, H · 2024
Closest in time.
Kangaroo: Lossless self-speculative decoding via double early exiting
Liu, F., Tang, Y., Liu, Z., Ni, Y., Han, K., and Wang, Y · 2024
Closest in time.
Specexec: Massively parallel speculative decoding for interactive llm inference on consumer devices
Svirschevski, R., May, A., Chen, Z., Chen, B., Jia, Z., and Ryabinin, M · 2024
Closest in time.
Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding, 2024
Xia, H., Yang, Z., Dong, Q., Wang, P., Li, Y., Ge, T., Liu, T., Li, W., and Sui, Z · 2024
Closest in time.
Yi, H., Lin, F., Li, H., Ning, P., Yu, X., and Xiao, R · 2024
Closest in time.
Recurrent drafter for fast speculative decoding in large language models
Zhang, A., Wang, C., Wang, Y., Zhang, X., and Cheng, Y · 2024
Closest in time.
Ouroboros: Speculative decoding with large model enhanced drafting
Zhao, W., Huang, Y., Han, X., Xiao, C., Liu, Z., and Sun, M · 2024
Closest in time.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2024
Closest in time.