Fetching the paper…
Reading the bibliography…
Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al · 2018
Earlier work this paper cites.
CodeSearchNet challenge: Evaluating the state of semantic code search
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Knowledge distillation as efficient pre-training: Faster convergence, higher data-efficiency, and better transferability
He, R., Sun, S., Yang, J., Bai, S., and Qi, X · 2022
Earlier work this paper cites.
URL https://huggingface.co/docs/transformers/en/index
Huggingface transformers, 2023 · 2023
Earlier work this paper cites.
URL https://huggingface.co/datasets/Aeala/ShareGPT_Vicuna_unfiltered
Sharegpt, 2023 · 2023
Cited alongside, same era.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2023
Cited alongside, same era.
gbharti/finance-alpaca, 2023
Bharti, G · 2023
Cited alongside, same era.
Medusa: Simple framework for accelerating llm generation with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., and Dao, T · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
llama.cpp, 2023
distilled bert topic, 2023
Luqmani, A. M · 2023
Closest in time.
Specinfer: Accelerating generative llm serving with speculative inference and token tree verification, 2023
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Closest in time.
Gpt-4 technical report
OpenAI, R · 2023
Closest in time.
Accelerating llm inference with staged speculative decoding
Spector, B. and Re, C · 2023
Closest in time.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2023
Closest in time.
Distillspec: Improving speculative decoding via knowledge distillation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gerganov, G · 2023
Cited alongside, same era.
Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2023
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
He, Z., Zhong, Z., Cai, T., Lee, J. D., and He, D · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J
Cited in the paper.
Cascade speculative drafting for even faster llm inference
Chen, Z., Yang, X., Lin, J., Sun, C., Huang, J., and Chang, K. C.-C
Cited in the paper.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al
Cited in the paper.
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2023
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Closest in time.
Bita: Bi-directional tuning for lossless acceleration in large language models
Lin, F., Yi, H., Li, H., Yang, Y., Yu, X., Lu, G., and Xiao, R · 2024
Closest in time.
Multi-candidate speculative decoding
Yang, S., Huang, S., Dai, X., and Chen, J · 2024
Closest in time.