Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks.
M. Stern, N. Shazeer, and J. Uszkoreit, “Blockwise parallel decoding for deep autoregressive models,” in NeurIPS , 2018, pp. 10 107–10 116
2018
Earlier work this paper cites.
G. Jawahar, B. Sagot, and D. Seddah, “What does BERT learn about the structure of language?” in ACL , 2019, pp. 3651–3657
2019
Earlier work this paper cites.
A. Rogers, O. Kovaleva, and A. Rumshisky, “A primer in bertology: What we know about how BERT works,” TACL , vol. 8, pp. 842–866, 2020
2020
Earlier work this paper cites.
Aeala, “Sharegpt_vicuna_unfiltered,” https://huggingface.co/datasets/Aeala/ShareGPT_Vicuna_unfiltered
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in ICML , vol. 202, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper, “Accelerating large language model decoding with speculative sampling,” 2023
2023
Cited alongside, same era.
H. Xia, T. Ge, P. Wang, S. Chen, F. Wei, and Z. Sui, “Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation,” in EMNLP , 2023, pp. 3909–3925
2023
Cited alongside, same era.
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer, “Speculative decoding with big little decoder,” in NeurIPS , 2023
2023
Cited alongside, same era.
Y. Zhou, K. Lyu, A. S. Rawat, A. K. Menon, A. Rostamizadeh, S. Kumar, J.-F. Kagy, and R. Agarwal, “Distillspec: Improving speculative decoding via knowledge distillation,” 2023
2023
Cited alongside, same era.
B. Spector and C. Re, “Accelerating llm inference with staged speculative decoding,” 2023
F. D. Keles, P. M. Wijewardena, and C. Hegde, “On the computational complexity of self-attention,” in ALT , vol. 201, 2023, pp. 597–619
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom, “Llama 2: Open foundation and fine-tuned chat models,” 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Z. Sun, A. T. Suresh, J. H. Ro, A. Beirami, H. Jain, and F. X. Yu, “Spectr: Fast speculative decoding via optimal transport,” in NeurIPS , 2023
2023
Cited alongside, same era.
Z. Chen, X. Yang, J. Lin, C. Sun, J. Huang, and K. C.-C. Chang, “Cascade speculative drafting for even faster llm inference,” 2023
2023
Cited alongside, same era.
A. Santilli, S. Severino, E. Postolache, V. Maiorca, M. Mancusi, R. Marin, and E. Rodolà, “Accelerating transformer inference for translation via parallel decoding,” in ACL , 2023, pp. 12 336–12 355
2023
Cited alongside, same era.
S. Yang, G. Lee, J. Cho, D. Papailiopoulos, and K. Lee, “Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding,” 2023
2023
Cited alongside, same era.
J. Zhang, J. Wang, H. Li, L. Shou, K. Chen, G. Chen, and S. Mehrotra, “Draft & \& verify: Lossless large language model acceleration via self-speculative decoding,” 2023
2023
Cited alongside, same era.
G. Monea, A. Joulin, and E. Grave, “Pass: Parallel speculative sampling,” 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
Q. Zheng, “Wku_nlp at semeval-2023 task 9: Translation augmented multilingual tweet intimacy analysis,” in ACL , 2023, pp. 1525–1530
2023
Later among the works it cites.
X. Liu, L. Hu, P. Bailis, I. Stoica, Z. Deng, A. Cheung, and H. Zhang, “Online speculative decoding,” 2023
2023
Later among the works it cites.
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao, “Medusa: Simple llm inference acceleration framework with multiple decoding heads,” arXiv preprint arXiv: 2401.10774 , 2024
2024
Closest in time.
2024
Closest in time.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, Z. Zhang, R. Y. Y. Wong, A. Zhu, L. Yang, X. Shi, C. Shi, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia, “Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification,” 2024
2024
Closest in time.
C. Hooper, S. Kim, H. Mohammadzadeh, H. Genc, K. Keutzer, A. Gholami, and S. Shao, “Speed: Speculative pipelined execution for efficient decoding,” 2024
2024
Closest in time.
Y. Fu, P. Bailis, I. Stoica, and H. Zhang, “Break the sequential dependency of llm inference using lookahead decoding,” 2024
2024
Closest in time.