Fetching the paper…
Reading the bibliography…
To improve the efficiency of distributed large language model (LLM) inference, various parallelization strategies, such as tensor and pipeline parallelism, have been proposed.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
PCI-SIG Releases PCIe 4.0, Version 1.0, 2017
PCI-SIG · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N · 2018
Earlier work this paper cites.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Ben-Nun, T. and Hoefler, T · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Earlier work this paper cites.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2021
Earlier work this paper cites.
{ \{ Zero-offload } \} : Democratizing { \{ billion-scale } \} model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Earlier work this paper cites.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Ryabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2022
Earlier work this paper cites.
Can foundation models wrangle your data?
Narayan, A., Chami, I., Orr, L., Arora, S., and Ré, C · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
Alpa: Automating inter-and { \{ Intra-Operator } \} parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Earlier work this paper cites.
Striped attention: Faster ring attention for causal transformers
Brandon, W., Nrusimha, A., Qian, K., Ankner, Z., Jin, T., Song, Z., and Ragan-Kelley, J · 2023
Earlier work this paper cites.
Hexgen: Generative inference of large language model over heterogeneous environment
Jiang, Y., Yan, R., Yao, X., Zhou, Y., Chen, B., and Yuan, B · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving
Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Ring attention with blockwise transformers for near-infinite context
Liu, H., Zaharia, M., and Abbeel, P · 2023
Earlier work this paper cites.
Towards efficient generative large language model serving: A survey from algorithms to systems
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Jin, H., Chen, T., and Jia, Z · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., et al · 2023
Cited alongside, same era.
Sharegpt vicuna unfiltered dataset
ShareGPT · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
P/d-serve: Serving disaggregated large language model at scale
Jin, Y., Wang, T., Lin, H., Song, M., Li, P., Ma, Y., Shan, Y., Yuan, Z., Li, C., Sun, Y., et al · 2024
Later among the works it cites.
Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models
Kamahori, K., Gu, Y., Zhu, K., and Kasikci, B · 2024
Later among the works it cites.
How bytedance scales offline inference with multi-modal llms to 200tb data
Kamsetty, A., Chen, H., and Xie, L · 2024
Later among the works it cites.
Infinite-llm: Efficient llm service for long context with distattention and distributed kvcache
Lin, B., Peng, T., Zhang, C., Sun, M., Li, L., Zhao, H., Xiao, W., Xu, Q., Qiu, X., Li, S., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Powerinfer: Fast large language model serving with a consumer-grade gpu
Song, Y., Mi, Z., Xie, H., and Chen, H · 2023
Cited alongside, same era.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., et al · 2023
Cited alongside, same era.
Efficiently programming large language models using sglang
Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Why not multiprocess pin memory in data loader?
AlbanD · 2024
Cited alongside, same era.
Amazon EC2 Instance Types, 2024
Amazon Web Services · 2024
Cited alongside, same era.
Snowflake llm inference: Optimizing gpu capacity for interactive workloads
Chan, V., Zhang, H., and Wang, F · 2024
Cited alongside, same era.
Liu, S., Biswal, A., Cheng, A., Mo, X., Cao, S., Gonzalez, J. E., Stoica, I., and Zaharia, M · 2024
Later among the works it cites.
Helix: Distributed serving of large language models via max-flow on heterogeneous gpus
Mei, Y., Zhuang, Y., Miao, X., Yang, J., Jia, Z., and Vinayak, R · 2024
Later among the works it cites.
Spotserve: Serving generative large language models on preemptible instances
Miao, X., Shi, C., Duan, J., Xi, X., Lin, D., Cui, B., and Jia, Z · 2024
Later among the works it cites.
Mlperf inference: Datacenter benchmark suite
MLCommons · 2024
Later among the works it cites.
Nvidia a100 pcie product brief, 2020
NVIDIA Corporation · 2024
Later among the works it cites.
NVIDIA NVLink: High-Speed GPU Interconnect, 2024
NVIDIA Corporation · 2024
Later among the works it cites.
Chatgpt (gpt-4), 2024
OpenAI · 2024
Later among the works it cites.
Instinfer: In-storage attention offloading for cost-effective long-context llm inference
Pan, X., Li, E., Li, Q., Liang, S., Shan, Y., Zhou, K., Luo, Y., Wang, X., and Zhang, J · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., and Bianchini, R · 2024
Later among the works it cites.
Mooncake: Kimi’s kvcache-centric architecture for llm serving
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Later among the works it cites.
Preble: Efficient distributed prompt scheduling for llm serving
Srivatsa, V., He, Z., Abhyankar, R., Li, D., and Zhang, Y · 2024
Later among the works it cites.
Performance update: Bringing vllm to the next level
vLLM Team · 2024
Later among the works it cites.
Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices
Xuanlei, Z., Jia, B., Zhou, H., Liu, Z., Cheng, S., and You, Y · 2024
Later among the works it cites.
Longvila: Scaling long-context visual language models for long videos
Xue, F., Chen, Y., Li, D., Hu, Q., Zhu, L., Li, X., Fang, Y., Tang, H., Yang, S., Liu, Z., et al · 2024
Later among the works it cites.
Batch llm inference on anyscale slashes aws bedrock costs by up to 6x, October 2024
Yu, C., Lee, S., Xu, R., Lin, W., Gorthy, P., and Liaw, R · 2024
Later among the works it cites.
Llm inference unveiled: Survey and roofline model insights
Yuan, Z., Shang, Y., Zhou, Y., Dong, Z., Xue, C., Wu, B., Li, Z., Gu, Q., Lee, Y. J., Yan, Y., et al · 2024
Later among the works it cites.
Llm-pq: Serving llm on heterogeneous clusters with phase-aware partition and adaptive quantization
Zhao, J., Wan, B., Peng, Y., Lin, H., and Wu, C · 2024
Later among the works it cites.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Later among the works it cites.
Nanoflow: Towards optimal large language model serving throughput
Zhu, K., Zhao, Y., Zhao, L., Zuo, G., Gu, Y., Xie, D., Gao, Y., Xu, Q., Tang, T., Ye, Z., et al · 2024
Later among the works it cites.