Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable performance, and organizations are racing to serve LLMs of varying sizes as endpoints for use-cases like chat, programming and search.
Salus: Fine-grained GPU sharing primitives for deep learning applications
Yu, P. and Chowdhury, M · 1902
Earlier work this paper cites.
Clipper: A Low-Latency online prediction serving system
Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., and Stoica, I · 2017
Earlier work this paper cites.
Tensorflow-serving: Flexible, high-performance ML serving
Olston, C., Fiedel, N., Gorovoy, K., Harmsen, J., Lao, L., Li, F., Rajashekhar, V., Ramesh, S., and Soyke, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E. Z., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Nexus: a gpu cluster engine for accelerating dnn-based video analysis
Shen, H., Chen, L., Jin, Y., Zhao, L., Kong, B., Philipose, M., Krishnamurthy, A., and Sundaram, R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
GSLICE: controlled spatial sharing of gpus for a scalable inference platform
Dhakal, A., Kulkarni, S. G., and Ramakrishnan, K. K · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
Gujarati, A., Karimi, R., Alzayat, S., Hao, W., Kaufmann, A., Vigfusson, Y., and Mace, J · 2020
Earlier work this paper cites.
AntMan: Dynamic scaling on GPU clusters for deep learning
Xiao, W., Ren, S., Li, Y., Zhang, Y., Hou, P., Li, Z., Feng, Y., Lin, W., and Jia, Y · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R. B., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N. D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M. S., Krishna, R., Kuditipudi, R., and et al · 2021
Earlier work this paper cites.
Zico: Efficient GPU memory sharing for concurrent DNN training
Lim, G., Ahn, J., Xiao, W., Kwon, Y., and Jeon, M · 2021
Cited alongside, same era.
Fastertransformer
NVIDIA · 2021
Cited alongside, same era.
INFaaS: Automated model-less inference serving
Romero, F., Li, Q., Yadwadkar, N. J., and Kozyrakis, C · 2021
Cited alongside, same era.
Serving DNN models with multi-instance gpus: A case of the reconfigurable machine scheduling problem
Tan, C., Li, Z., Zhang, J., Cao, Y., Qi, S., Liu, Z., Zhu, Y., and Guo, C · 2021
Cited alongside, same era.
Wavelet: Efficient DNN training with tick-tock scheduling
Wang, G., Wang, K., Jiang, K., Li, X., and Stoica, I · 2021
Cited alongside, same era.
Lightseq: A high performance inference library for transformers
Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L · 2021
Text generation inference
Huggingface · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
Alpaserve: Statistical multiplexing with model parallelism for deep learning serving
Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification, 2023
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
Sharegpt
ShareGPT-Team · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., and He, Y · 2022
Cited alongside, same era.
Serving heterogeneous machine learning models on Multi-GPU servers with Spatio-Temporal sharing
Choi, S., Lee, S., Kim, Y., Park, J., Kwon, Y., and Huh, J · 2022
Cited alongside, same era.
Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences
Han, M., Zhang, H., Chen, R., and Chen, H · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Cited alongside, same era.
Alpa: Automating inter- and intra-operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2023
Cited alongside, same era.
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
SHEPHERD: Serving DNNs in the wild
Zhang, H., Tang, Y., Khandelwal, A., and Stoica, I · 2023
Later among the works it cites.
Muxflow: Efficient and safe GPU sharing in large-scale production deep learning clusters
Zhao, Y., Liu, X., Liu, S., Li, X., Zhu, Y., Huang, G., Liu, X., and Jin, X · 2023
Later among the works it cites.
Spotserve: Serving generative large language models on preemptible instances
Miao, X., Shi, C., Duan, J., Xi, X., Lin, D., Cui, B., and Jia, Z · 2024
Closest in time.
Fairness in serving large language models
Sheng, Y., Cao, S., Li, D., Zhu, B., Li, Z., Zhuo, D., Gonzalez, J. E., and Stoica, I · 2024
Closest in time.