Fetching the paper…
Reading the bibliography…
Optimizing the deployment of Large language models (LLMs) is expensive today since it requires experimentally running an application workload against an LLM implementation while exploring large configuration space formed by system knobs such as parallelization strategies, batching techniques, and scheduling policies.
Roofline: An insightful visual performance model for multicore architectures
Williams, S., Waterman, A., and Patterson, D · 2009
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning, 2014
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Gandiva: Introspective cluster scheduling for deep learning
Xiao, W., Bhardwaj, R., Ramjee, R., Sivathanu, M., Kwatra, N., Han, Z., Patel, P., Peng, X., Zhao, H., Zhang, Q., Yang, F., and Zhou, L · 2018
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Astra: Exploiting predictability to optimize deep learning
Sivathanu, M., Chugh, T., Singapuram, S. S., and Zhou, L · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Daydream: Accurately estimating the efficacy of optimizations for DNN training
Zhu, H., Phanishayee, A., and Pekhimenko, G · 2020
Earlier work this paper cites.
Terapipe: Token-level pipeline parallelism for training large-scale language models, 2021
Li, Z., Zhuang, S., Guo, S., Zhuo, D., Zhang, H., Song, D., and Stoica, I · 2021
Earlier work this paper cites.
Habitat: A runtime-based computational performance predictor for deep neural network training
Yu, G. X., Gao, Y., Golikov, P., and Pekhimenko, G · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Building a performance model for deep learning recommendation model training on gpus
Lin, Z., Feng, L., Ardestani, E. K., Lee, J., Lundell, J., Kim, C., Kejariwal, A., and Owens, J. D · 2022
Cited alongside, same era.
Efficiently scaling transformer inference, 2022
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Cited alongside, same era.
https://github.com/ModelTC/lightllm , 2023
LightLLM: A python-based large language model inference and serving framework · 2023
Cited alongside, same era.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills, 2023
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Cited alongside, same era.
Discourse centric evaluation of machine translation with a densely annotated parallel corpus
Jiang, Y. E., Liu, T., Ma, S., Zhang, D., Cotterell, R., and Sachan, M · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T · 2023
Later among the works it cites.
The inference cost of search disruption – large language model cost analysis, 2023
Patel, D. and Ahmed, A · 2023
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Patel, P., Choukse, E., Zhang, C., Goiri, Í., Shah, A., Maleki, S., and Bianchini, R · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Cited alongside, same era.
Flash-decoding for long-context inference, 2023
Dao, T., Haziza, D., Massa, F., and Sizov, G · 2023
Cited alongside, same era.
Proteus: Simulating the performance of distributed DNN training
Duan, J., Li, X., Xu, P., Zhang, X., Yan, S., Liang, Y., and Lin, D · 2023
Cited alongside, same era.
Speed: Speculative pipelined execution for efficient decoding, 2023
Hooper, C., Kim, S., Mohammadzadeh, H., Genc, H., Keutzer, K., Gholami, A., and Shao, S · 2023
Cited alongside, same era.
https://arxiv.org/
arxiv.org e-print archive
Cited in the paper.
https://docs.nvidia.com/cuda/cupti/index.html
Cupti: Cuda toolkit documentation
Cited in the paper.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E. P., Gonzalez, J. E., Stoica, I., and Zhang, H · 2023
Later among the works it cites.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.
A discourse-aware attention model for abstractive summarization of long documents
Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N · 2097
Closest in time.