Fetching the paper…
Reading the bibliography…
Serving generative inference of the large language model is a crucial component of contemporary AI applications.
Bulk synchronous parallel computing—a paradigm for transportable software
Cheatham, T., Fahmy, A., Stefanescu, D., and Valiant, L · 1996
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Pipedream: generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Distributed deep learning in open collaborations
Diskin, M., Bukhtiyarov, A., Ryabinin, M., Saulnier, L., Sinitsin, A., Popov, D., Pyrkin, D. V., Kashirin, M., Borzunov, A., Villanova del Moral, A., et al · 2021
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Earlier work this paper cites.
From cloud computing to sky computing
Stoica, I. and Shenker, S · 2021
Earlier work this paper cites.
Optimizing inference serving on serverless platforms
Ali, A., Pinciroli, R., Yan, F., and Smirni, E · 2022
Earlier work this paper cites.
Varuna: scalable, low-cost training of massive deep learning models
Athlur, S., Saran, N., Sivathanu, M., Ramjee, R., and Kwatra, N · 2022
Earlier work this paper cites.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Ryabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Hydrozoa: Dynamic hybrid-parallel dnn training on serverless containers
Guo, R., Guo, V., Kim, A., Hildred, J., and Daudjee, K · 2022
Earlier work this paper cites.
Hugging face accelerate
HuggingFace · 2022
Earlier work this paper cites.
Kwon, S. J., Kim, J., Bae, J., Yoo, K. M., Kim, J.-H., Park, B., Kim, B., Ha, J.-W., Sung, N., and Lee, D · 2022
Cited alongside, same era.
Fastertransformer
NVIDIA · 2022
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Closest in time.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Closest in time.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving
Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., et al · 2023
Closest in time.
A modular network stack, 2023
LibP2P · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Closest in time.
Deja vu: Contextual sparsity for efficient llms at inference time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decentralized training of foundation models in heterogeneous environments
Yuan, B., He, Y., Davis, J., Zhang, T., Dao, T., Chen, B., Liang, P. S., Re, C., and Zhang, C · 2022
Cited alongside, same era.
Sakshi: Decentralized ai platforms
Bhat, S., Chen, C., Cheng, Z., Fang, Z., Hebbar, A., Kannan, S., Rana, R., Sheng, P., Tyagi, H., Viswanath, P., et al · 2023
Cited alongside, same era.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Riabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Medusa: Simple framework for accelerating llm generation with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., and Dao, T · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Closest in time.
Chatbot arena conversations
Lmsys · 2023
Closest in time.
Swarm parallelism: Training large models can be surprisingly communication-efficient
Ryabinin, M., Dettmers, T., Diskin, M., and Borzunov, A · 2023
Closest in time.
Accelerating llm inference with staged speculative decoding
Spector, B. F. and Re, C · 2023
Closest in time.
Bamboo: Making preemptible instances resilient for affordable training of large { \{ DNNs } \}
Thorpe, J., Zhao, P., Eyolfson, J., Qiao, Y., Jia, Z., Zhang, M., Netravali, R., and Xu, G. H · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
{ \{ SkyPilot } \} : An intercloud broker for sky computing
Yang, Z., Wu, Z., Luo, M., Chiang, W.-L., Bhardwaj, R., Kwon, W., Zhuang, S., Luan, F. S., Mittal, G., Shenker, S., et al · 2023
Closest in time.
Open Compute Framework: Peer-to-Peer Task Queue for Foundation Model Inference Serving, September 2023
Yao, X · 2023
Closest in time.
Efficient distributed transaction processing in heterogeneous networks
Zhang, Q., Li, J., Zhao, H., Xu, Q., Lu, W., Xiao, J., Han, F., Yang, C., and Du, X · 2023
Closest in time.