Fetching the paper…
Reading the bibliography…
Online LLM inference powers many exciting applications such as intelligent chatbots and autonomous agents.
Attention is All You Need
Vaswani, A · 2017
Earlier work this paper cites.
Ray: A distributed framework for emerging { \{ AI } \} applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., et al · 2018
Earlier work this paper cites.
Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations
Tillet, P., Kung, H.-T., and Cox, D · 2019
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Earlier work this paper cites.
GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Kwon, S. J., Kim, J., Bae, J., Yoo, K. M., Kim, J.-H., Park, B., Kim, B., Ha, J.-W., Sung, N., and Lee, D · 2022
Earlier work this paper cites.
NVIDIA A100 Tensor Core GPU
NVIDIA · 2022
Earlier work this paper cites.
ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-Based Generative Models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
Flash-Decoding for Long-Context Inference
Dao, T., Haziza, D., Massa, F., and Sizov, G · 2023
Cited alongside, same era.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
FlashDecoding++: Faster Large Language Model Inference on GPUs
Hong, K., Dai, G., Xu, J., Mao, Q., Li, X., Liu, J., Chen, K., Dong, Y., and Wang, Y · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Azure LLM inference trace 2023
Microsoft · 2023
Cited alongside, same era.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
He, J. and Zhai, J · 2024
Closest in time.
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Hu, C., Huang, H., Xu, L., Chen, X., Xu, J., Chen, S., Feng, H., Wang, C., Wang, S., Bao, Y., et al · 2024
Closest in time.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Closest in time.
NVIDIA H100 Tensor Core GPU
NVIDIA · 2024
Closest in time.
Instinfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
Pan, X., Li, E., Li, Q., Liang, S., Shan, Y., Zhou, K., Luo, Y., Wang, X., and Zhang, J · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Fast Distributed Inference Serving for Large Language Models
Wu, B., Zhong, Y., Zhang, Z., Huang, G., Liu, X., and Jin, X · 2023
Cited alongside, same era.
SGLang: Efficient Execution of Structured Language Model Programs
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Efficient and Economic Large Language Model Inference with Attention Offloading
Chen, S., Lin, Y., Zhang, M., and Wu, Y · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Closest in time.
Splitwise: Efficient Generative LLM Inference Using Phase Splitting
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., and Bianchini, R · 2024
Closest in time.
PowerInfer: Fast Large Language Model Serving with a Consumer-Grade GPU
Song, Y., Mi, Z., Xie, H., and Chen, H · 2024
Closest in time.
DejaVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
Strati, F., Mcallister, S., Phanishayee, A., Tarnawski, J., and Klimovic, A · 2024
Closest in time.
TwinPilots: A New Computing Paradigm for GPU-CPU Parallel LLM Inference
Yu, C., Wang, T., Shao, Z., Zhu, L., Zhou, X., and Jiang, S · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.
NanoFlow: Towards Optimal Large Language Model Serving Throughput
Zhu, K., Zhao, Y., Zhao, L., Zuo, G., Gu, Y., Xie, D., Gao, Y., Xu, Q., Tang, T., Ye, Z., et al · 2024
Closest in time.