Fetching the paper…
Reading the bibliography…
Large language model (LLM)-based inference workloads increasingly dominate data center costs and resource utilization.
Data center evolution: A tutorial on state of the art, issues, and challenges
Krishna Kant · 2009
Earlier work this paper cites.
Kernel fusion: An effective method for better power efficiency on multithreaded gpu
Guibin Wang, YiSong Lin, and Wei Yi · 2010
Earlier work this paper cites.
Automatic fusions of cuda-gpu kernels for parallel map
Jan Fousek, Jiři Filipovič, and Matuš Madzin · 2011
Earlier work this paper cites.
Scalable kernel fusion for memory-bound gpu applications
Mohamed Wahib and Naoya Maruyama · 2014
Earlier work this paper cites.
The features, hardware, and architectures of data center networks: A survey
Tao Chen, Xiaofeng Gao, and Guihai Chen · 2016
Earlier work this paper cites.
Heterogeneous integration for performance and scaling
Subramanian S Iyer · 2016
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Getting started with cuda graphs
Alan Gray · 2019
Earlier work this paper cites.
Taso: optimizing deep learning computation with automatic generation of graph substitutions
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken · 2019
Earlier work this paper cites.
Nexus: A gpu cluster engine for accelerating dnn-based video analysis
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, Hsiang-Tsung Kung, and David Cox · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Astra-sim: Enabling sw/hw co-design exploration for distributed dl training platforms
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna · 2020
Earlier work this paper cites.
Daydream: Accurately estimating the efficacy of optimizations for { \{ DNN } \} training
Hongyu Zhu, Amar Phanishayee, and Gennady Pekhimenko · 2020
Earlier work this paper cites.
Lazy batching: An sla-aware batching system for cloud machine learning inference
Yujeong Choi, Yunseong Kim, and Minsoo Rhu · 2021
Earlier work this paper cites.
Introduction to pytorch
Nikhil Ketkar, Jojo Moolayil, Nikhil Ketkar, and Jojo Moolayil · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
dpro: A generic performance diagnosis and optimization toolkit for expediting distributed dnn training
Hanpeng Hu, Chenyu Jiang, Yuchen Zhong, Yanghua Peng, Chuan Wu, Yibo Zhu, Haibin Lin, and Chuanxiong Guo · 2022
Earlier work this paper cites.
Building a performance model for deep learning recommendation model training on gpus
Zhongyi Lin, Louis Feng, Ehsan K Ardestani, Jaewon Lee, John Lundell, Changkyu Kim, Arun Kejariwal, and John D Owens · 2022
Cited alongside, same era.
Operationalizing and implementing pretrained, large artificial intelligence linguistic models in the us health care system: outlook of generative pretrained transformer 3 (gpt-3) as a service model
Emre Sezgin, Joseph Sirrianni, and Simon L Linwood · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
https://research.ibm.com/blog/AI-inference-explained, 2023
What is ai inferencing? · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
https://developer.nvidia.com/cupti
Nvidia cuda profiling tools interface (cupti) - cuda toolkit · 2024
Later among the works it cites.
https://developer.nvidia.com/nsight-compute
Nvidia nsight compute · 2024
Later among the works it cites.
https://developer.nvidia.com/nsight-systems
Nvidia nsight sytems · 2024
Later among the works it cites.
Luigi Fusco, Mikhail Khalilov, Marcin Chrapek, Giridhar Chukkapalli, Thomas Schulthess, and Torsten Hoefler · 2024
Later among the works it cites.
Generative-ai in e-commerce: Use-cases and implementations
Shervin Ghaffari, Behnam Yousefimehr, and Mehdi Ghatee · 2024
Later among the works it cites.
Preliminary performance evaluation of grace-hopper gh200
Toshihiro Hanawa, Kengo Nakajima, Yohei Miki, Takashi Shimokawabe, Kazuya Yamazaki, Shinji Sumimoto, Osamu Tatebe, Taisuke Boku, Daisuke Takahashi, Akira Nukada, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trends in ai inference energy consumption: Beyond the performance-vs-parameter laws of deep learning
Radosvet Desislavov, Fernando Martínez-Plumed, and José Hernández-Orallo · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2023
Cited alongside, same era.
The framework tax: Disparities between inference efficiency in nlp research and deployment
Jared Fernandez, Jacob Kahn, Clara Na, Yonatan Bisk, and Emma Strubell · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu · 2023
Cited alongside, same era.
Performance Modeling and Optimization for Machine Learning Workloads
Zhongyi Lin · 2023
Cited alongside, same era.
Polca: Power oversubscription in llm cloud providers
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Brijesh Warrier, Nithish Mahalingam, and Ricardo Bianchini · 2023
Cited alongside, same era.
Later among the works it cites.
Human latency conversational turns for spoken avatar systems
Derek Jacoby, Tianyi Zhang, Aanchan Mohan, and Yvonne Coady · 2024
Later among the works it cites.
The genai is out of the bottle: generative artificial intelligence from a business model innovation perspective
Dominik K Kanbach, Louisa Heiduk, Georg Blueher, Maximilian Schreiter, and Alexander Lahmann · 2024
Later among the works it cites.
Automatic blas offloading on unified memory architecture: A study on nvidia grace-hopper
Junjie Li, Yinzhi Wang, Xiao Liang, and Hang Liu · 2024
Later among the works it cites.
Fine-grained trace-driven performance modeling and simulation for large-scale ml training
Mingyu Liang, Hiwot Tadese Kassa, Wenyin Fu, Brian Coutinho, Louis Feng, and Christina Delimitrou · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Later among the works it cites.
First impressions of the nvidia grace cpu superchip and nvidia grace hopper superchip for scientific workloads
Nikolay A Simakov, Matthew D Jones, Thomas R Furlani, Eva Siegmann, and Robert J Harrison · 2024
Later among the works it cites.
11.1 amd instincttm mi300 series modular chiplet package–hpc and ai accelerator for exa-class systems
Alan Smith, Eric Chapman, Chintan Patel, Raja Swaminathan, John Wuu, Tyrone Huang, Wonjun Jung, Alexander Kaganov, Hugh McIntyre, and Ramon Mangaser · 2024
Later among the works it cites.
Towards greener llms: Bringing energy-efficiency to the forefront of llm inference
Jovan Stojkovic, Esha Choukse, Chaojie Zhang, Inigo Goiri, and Josep Torrellas · 2024
Later among the works it cites.
Introduction to torch.compile
William Wen · 2024
Later among the works it cites.
Llm inference unveiled: Survey and roofline model insights
Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, et al · 2024
Later among the works it cites.
https://mlcommons.org/benchmarks/inference-datacenter/
Mlperf inference: Datacenter · 2025
Closest in time.