Fetching the paper…
Reading the bibliography…
Each request in LLM inference goes through two phases: compute-bound prefill and memory-bandwidth-bound decode.
A study of Persistent Threads style GPU programming for GPGPU workloads. In 2012 Innovative Parallel Computing (InPar) . 1–14
Kshitij Gupta, Jeff A. Stuart, and John D. Owens. 2012 · 2012
Earlier work this paper cites.
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation. In 2012 45th Annual IEEE/ACM International Symposium on Microarchitecture . 107–118
Haicheng Wu, Gregory Diamos, Srihari Cadambi, and Sudhakar Yalamanchili. 2012 · 2012
Earlier work this paper cites.
Improving GPGPU concurrency with elastic kernels. In Proceedings of the Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems (Houston, Texas, USA) (ASPLOS ’13) . Association for Computing Machinery, New York, NY, USA, 407–418
Sreepathi Pai, Matthew J. Thazhuthaveetil, and R. Govindarajan. 2013 · 2013
Earlier work this paper cites.
Design and evaluation of the gemtc framework for GPU-enabled many-task computing. In Proceedings of the 23rd International Symposium on High-Performance Parallel and Distributed Computing (Vancouver, BC, Canada) (HPDC ’14) . Association for Computing Machinery, New York, NY, USA, 153–164
Scott J. Krieder, Justin M. Wozniak, Timothy Armstrong, Michael Wilde, Daniel S. Katz, Benjamin Grimmer, Ian T. Foster, and Ioan Raicu. 2014 · 2014
Earlier work this paper cites.
Efficient GPU Spatial-Temporal Multitasking
Yun Liang, Huynh Phung Huynh, Kyle Rupnow, Rick Siow Mong Goh, and Deming Chen. 2015 · 2014
Earlier work this paper cites.
Scalable Kernel Fusion for Memory-Bound GPU Applications. In SC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 191–202
Mohamed Wahib and Naoya Maruyama. 2014 · 2014
Earlier work this paper cites.
Enabling and Exploiting Flexible Task Assignment on GPU through SM-Centric Program Transformations. In Proceedings of the 29th ACM on International Conference on Supercomputing (Newport Beach, California, USA) (ICS ’15) . Association for Computing Machinery, New York, NY, USA, 119–130
Bo Wu, Guoyang Chen, Dong Li, Xipeng Shen, and Jeffrey Vetter. 2015 · 2015
Earlier work this paper cites.
Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharing. In 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA) . 358–369
Zhenning Wang, Jun Yang, Rami Melhem, Bruce Childers, Youtao Zhang, and Minyi Guo. 2016 · 2016
Earlier work this paper cites.
Programming Tensor Cores in CUDA 9
Jeremy Appleyard and Scott Yokim. 2017 · 2017
Earlier work this paper cites.
Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks. In Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (Austin, Texas, USA) (PPoPP ’17) . Association for Computing Machinery, New York, NY, USA, 221–234
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, and Timothy G. Rogers. 2017 · 2017
Earlier work this paper cites.
FlashAttention
2022 · 2022
Earlier work this paper cites.
FLASHATTENTION: fast and memory-efficient exact attention with IO-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’22) . Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland) (ASPLOS ’22) . Association for Computing Machinery, New York, NY, USA, 402–416
Abhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet, Saeed Maleki, Youshan Miao, Madanlal Musuvathi, Todd Mytkowicz, and Olli Saarikivi. 2022 · 2022
Earlier work this paper cites.
Automatic Horizontal Fusion for GPU Kernels. In 2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . 14–27
Ao Li, Bojian Zheng, Gennady Pekhimenko, and Fan Long. 2022 · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . USENIX Association, Carlsbad, CA, 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Earlier work this paper cites.
ISPA: Exploiting Intra-SM Parallelism in GPUs via Fine-Grained Resource Management
Han Zhao, Weihao Cui, Quan Chen, and Minyi Guo. 2023 · 2022
Earlier work this paper cites.
TensorRT-LLM: A TensorRT Toolbox for Optimized Large Language Model Inference
2023 · 2023
Earlier work this paper cites.
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023 · 2023
Earlier work this paper cites.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 4895–4901
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023 · 2023
Cited alongside, same era.
Flash-Decoding for long-context inference
Tri Dao, Daniel Haziza, Francisco Massa, and Grigory Sizov. 2023 · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (Koblenz, Germany) (SOSP ’23) . Association for Computing Machinery, New York, NY, USA, 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Stream-K: Work-Centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU. In Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (Montreal, QC, Canada) (PPoPP ’23) . Association for Computing Machinery, New York, NY, USA, 429–431
FlashDecoding++: Faster Large Language Model Inference with Asynchronization, Flat GEMM Optimization, and Heuristics. In Proceedings of Machine Learning and Systems , P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.), Vol. 6. 148–161
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, kangdi chen, Yuhan Dong, and Yu Wang. 2024 · 2024
Closest in time.
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan. 2024 · 2024
Closest in time.
Toward Efficient Inference for Mixture of Experts. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Haiyang Huang, Newsha Ardalani, Anna Sun, Liu Ke, Shruti Bhosale, Hsien-Hsin S. Lee, Carole-Jean Wu, and Benjamin Lee. 2024 · 2024
Closest in time.
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels. In Proceedings of the 2024 IEEE/ACM International Symposium on Code Generation and Optimization (Edinburgh, United Kingdom) (CGO ’24) . IEEE Press, 93–105
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland, and John D. Owens. 2023 · 2023
Cited alongside, same era.
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (Vancouver, BC, Canada) (ASPLOS 2023) . Association for Computing Machinery, New York, NY, USA, 93–106
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, Sameer Kumar, Tongfei Guo, Yuanzhong Xu, and Zongwei Zhou. 2022 · 2023
Cited alongside, same era.
Fast Distributed Inference Serving for Large Language Models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023 · 2023
Cited alongside, same era.
ccdv/arxiv-summarization
2024 · 2024
Cited alongside, same era.
CUDA C Programming Guide – Hardware Implementation
2024 · 2024
Cited alongside, same era.
NVIDIA/cutlass: CUDA Templates for Linear Algebra Subroutines
2024 · 2024
Cited alongside, same era.
The State of AI Infrastructure at Scale 2024
2024b · 2024
Cited alongside, same era.
vLLM: Easy, fast, and cheap LLM serving for everyone
2024 · 2024
Cited alongside, same era.
Yi-6B-200K
2024 · 2024
Cited alongside, same era.
Abhinav Jangda, Saeed Maleki, Maryam Mehri Dehnavi, Madan Musuvathi, and Olli Saarikivi. 2024 · 2024
Closest in time.
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
Hao Kang, Srikant Bharadwaj, James Hensman, Tushar Krishna, Victor Ruhle, and Saravan Rajmohan. 2024 · 2024
Closest in time.
Splitwise: Efficient Generative LLM Inference Using Phase Splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . 118–132
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024 · 2024
Closest in time.
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
Rya Sanovar, Srikant Bharadwaj, Renee St. Amant, Victor Rühle, and Saravan Rajmohan. 2024 · 2024
Closest in time.
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. 2024 · 2024
Closest in time.
Fairness in Serving Large Language Models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 965–988
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Closest in time.
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (Austin, TX, USA) (SOSP ’24) . Association for Computing Machinery, New York, NY, USA, 590–606
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. 2024 · 2024
Closest in time.
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2024 · 2024
Closest in time.
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (Austin, TX, USA) (SOSP ’24) . Association for Computing Machinery, New York, NY, USA, 640–654
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024 · 2024
Closest in time.
SGLang: Efficient Execution of Structured Language Model Programs. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024 · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 193–210
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.
NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Yufei Gao, Qinyu Xu, Tian Tang, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci. 2024 · 2024
Closest in time.
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (Rotterdam, Netherlands) (ASPLOS ’25) . Association for Computing Machinery, New York, NY, USA, 1133–1150
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2025 · 2025
Closest in time.
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze. 2025 · 2025
Closest in time.