Fetching the paper…
Reading the bibliography…
Speculative decoding has been shown as an effective way to accelerate Large Language Model (LLM) inference by using a Small Speculative Model (SSM) to generate candidate tokens in a so-called speculation phase, which are subsequently verified by the LLM in a verification phase.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 389–402
2021
Earlier work this paper cites.
B. Fu, F. Chen, P. Li, and D. Zeng, “Tcb: Accelerating transformer inference services with request concatenation,” in Proceedings of the 51st International Conference on Parallel Processing , 2022, pp. 1–11
2022
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for transformer–based generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Wang, K. Chen, H. Tan, and K. Guo, “Tabi: An efficient multi-level inference system for large language models,” in Proceedings of the Eighteenth European Conference on Computer Systems , 2023, pp. 233–248
2023
Earlier work this paper cites.
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Re et al. , “Deja vu: Contextual sparsity for efficient llms at inference time,” in International Conference on Machine Learning . PMLR, 2023, pp. 22 137–22 176
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
Y. Zhai, C. Jiang, L. Wang, X. Jia, S. Zhang, Z. Chen, X. Liu, and Y. Zhu, “Bytetransformer: A high-performance transformer boosted for variable-length inputs,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2023, pp. 344–355
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Rashad, “Chatgpt-prompts,” 2023. [Online]. Available: https://huggingface.co/datasets/MohamedRashad/ChatGPT-prompts
2023
Cited alongside, same era.
A. Palla, “Chatbot instruction prompts,” 2023. [Online]. Available: https://huggingface.co/datasets/alespalla/chatbot_instruction_prompts
2023
Cited alongside, same era.
HuggingFace, “Large language model text generation inference,” 2023. [Online]. Available: https://github.com/huggingface/text-generation-inference
2023
Cited alongside, same era.
Y. Fu, P. Bailis, I. Stoica, and H. Zhang, “Breaking the sequential dependency of llm inference using lookahead decoding,” November 2023. [Online]. Available: https://lmsys.org/blog/2023-11-21-lookahead-decoding/
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Sun, Z. Huang, H. Zhao, W. Xiao, X. Zhang, Y. Li, and W. Lin, “Llumnix: Dynamic scheduling for large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, 2024, pp. 173–191
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
H. Wen, Y. Li, G. Liu, S. Zhao, T. Yu, T. J.-J. Li, S. Jiang, Y. Liu, Y. Zhang, and Y. Liu, “Autodroid: Llm-powered task automation in android,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Networking , 2024, pp. 543–557
2024
Cited alongside, same era.
A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, and R. Ramjee, “Taming throughput-latency tradeoff in llm inference with sarathi-serve,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, 2024, pp. 117–134
2024
Cited alongside, same era.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, 2024, pp. 193–210
2024
Cited alongside, same era.
B. Sun, Z. Huang, H. Zhao, W. Xiao, X. Zhang, Y. Li, and W. Lin, “Llumnix: Dynamic scheduling for large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, 2024, pp. 173–191
2024
Cited alongside, same era.
W. Lee, J. Lee, J. Seo, and J. Sim, “Infinigen: Efficient generative inference of large language models with dynamic KV cache management,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, 2024, pp. 155–172
2024
Cited alongside, same era.
F. Chen, P. Li, S. Pan, L. Zhong, and J. Deng, “Giant could be tiny: Efficient inference of giant models on resource-constrained uavs,” IEEE Internet of Things Journal , 2024
2024
Cited alongside, same era.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, Z. Zhang, R. Y. Y. Wong, A. Zhu, L. Yang, X. Shi et al. , “Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 932–949
2024
Cited alongside, same era.
C. Li, Z. Zhou, S. Zheng, J. Zhang, Y. Liang, and G. Sun, “Specpim: Accelerating speculative inference on pim-enabled system via architecture-dataflow co-exploration,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 950–965
2024
Later among the works it cites.
B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, J. Deng, X. Yang, Z. Yu, and P. Zuo, “Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) , 2024, pp. 111–126
2024
Later among the works it cites.
“Technical report,” 2024. [Online]. Available: https://www.dropbox.com/scl/fi/1wl4pdw8z69e0taflcyuz/Technical_Report_SPIN.pdf?rlkey=815bdk0xg8fwrais3isriuygn&st=4n0d3a77&dl=0
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer, “Speculative decoding with big little decoder,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.