Fetching the paper…
Reading the bibliography…
The rapid evolution and widespread adoption of generative large language models (LLMs) have made them a pivotal workload in various applications.
T. Komoda, S. Hayashi, T. Nakada, S. Miwa, and H. Nakamura, “Power capping of cpu-gpu heterogeneous systems through coordinating dvfs and task mapping,” in Proceedings of the IEEE 31st International Conference on Computer Design (ICCD ’13) , 2013
2013
Earlier work this paper cites.
D. Lo, L. Cheng, R. Govindaraju, L. A. Barroso, and C. Kozyrakis, “Towards energy proportionality for large-scale latency-critical workloads,” in Proceedings of the ACM/IEEE 41st International Symposium on Computer Architecture (ISCA ’14) , 2014
2014
Earlier work this paper cites.
C.-H. Hsu, Y. Zhang, M. A. Laurenzano, D. Meisner, T. Wenisch, J. Mars, L. Tang, and R. G. Dreslinski, “Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting,” in Proceedings of the IEEE 21st International Symposium on High Performance Computer Architecture (HPCA ’15) , 2015
2015
Earlier work this paper cites.
H. Kasture, D. B. Bartolini, N. Beckmann, and D. Sanchez, “Rubik: Fast analytical power management for latency-critical systems,” in Proceedings of the 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’15) , 2015
2015
Earlier work this paper cites.
M. E. Haque, Y. He, S. Elnikety, T. D. Nguyen, R. Bianchini, and K. S. McKinley, “Exploiting Heterogeneity for Tail Latency and Energy Efficiency,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’17) , 2017
2017
Earlier work this paper cites.
C. Delimitrou and C. Kozyrakis, “Amdahl’s Law for Tail Latency,” Commun. ACM , vol. 61, no. 8, jul 2018
2018
Earlier work this paper cites.
C.-H. Hsu, Q. Deng, J. Mars, and L. Tang, “SmoothOperator: Reducing Power Fragmentation and Improving Power Utilization in Large-Scale Datacenters,” in Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’18) , 2018
2018
Earlier work this paper cites.
S. Chen, C. Delimitrou, and J. F. Martínez, “PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’19) , 2019
2019
Earlier work this paper cites.
M. Halpern, B. Boroujerdian, T. Mummert, E. Duesterwald, and V. Janapa Reddi, “One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-Off in Machine Learning Cloud Service APIs via Tolerance Tiers,” in Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS ’19) , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
Z. Tang, Y. Wang, Q. Wang, and X. Chu, “The Impact of GPU DVFS on the Energy and Performance of Deep Learning: an Empirical Study,” in Proceedings of the Tenth ACM International Conference on Future Energy Systems (e-Energy ’19) , 2019
2019
Earlier work this paper cites.
Y. G. Kim and C.-J. Wu, “AutoScale: Energy Efficiency Optimization for Stochastic Edge Inference Using Reinforcement Learning,” in Proceedings of the 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’20) , 2020
2020
Earlier work this paper cites.
N. Kulkarni, G. Gonzalez-Pumariega, A. Khurana, C. A. Shoemaker, C. Delimitrou, and D. H. Albonesi, “CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores,” in Proceedings of the 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’20) , 2020
2020
Earlier work this paper cites.
S. Li, X. Wang, X. Zhang, V. Kontorinis, S. Kodakara, D. Lo, and P. Ranganathan, “Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at Scale,” in Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’20) , 2020
2020
Earlier work this paper cites.
R. Nishtala, V. Petrucci, P. Carpenter, and M. Sjalander, “Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud Services,” in Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA ’20) , 2020
2020
Earlier work this paper cites.
T. Patel and D. Tiwari, “CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale Computers,” in Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA ’20) , 2020
2020
Earlier work this paper cites.
C. Wan, M. Santriaji, E. Rogers, H. Hoffmann, M. Maire, and S. Lu, “ALERT: Accurate learning for energy and timeliness,” in Proceedings of the USENIX Annual Technical Conference (USENIX ATC ’20) , 2020
2020
Earlier work this paper cites.
W. Xiao, S. Ren, Y. Li, Y. Zhang, P. Hou, Z. Li, Y. Feng, W. Lin, and Y. Jia, “AntMan: Dynamic scaling on GPU clusters for deep learning,” in Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’20) , 2020
2020
Earlier work this paper cites.
L. Zhou, L. N. Bhuyan, and K. K. Ramakrishnan, “Gemini: Learning to Manage CPU Power for Latency-Critical Search Engines,” in Proceedings of the 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’20) , 2020
2020
Earlier work this paper cites.
P. Zou, A. Li, K. Barker, and R. Ge, “Indicator-Directed Dynamic Power Management for Iterative Workloads on GPU-Accelerated Systems,” in Proceedings of the 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID ’20) , 2020
2020
Earlier work this paper cites.
A. G. Kumbhare, R. Azimi, I. Manousakis, A. Bonde, F. Frujeri, N. Mahalingam, P. A. Misra, S. A. Javadi, B. Schroeder, M. Fontoura, and R. Bianchini, “Prediction-Based Power Oversubscription in Cloud Platforms,” in Proceedings of the USENIX Annual Technical Conference (USENIX ATC ’21) , 2021
2021
Earlier work this paper cites.
A. Mirhosseini and T. Wenisch, “ μ \mu Steal: A Theory-Backed Framework for Preemptive Work and Resource Stealing in Mixed-Criticality Microservices,” in Proceedings of the ACM International Conference on Supercomputing (ICS ’21) , 2021
2021
Earlier work this paper cites.
S. M. Nabavinejad, S. Reda, and M. Ebrahimi, “BatchSizer: Power-Performance Trade-off for DNN Inference,” in Proceedings of the 26th Asia and South Pacific Design Automation Conference (ASP-DAC ’21) , 2021
2021
Earlier work this paper cites.
F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “INFaaS: Automated Model-less Inference Serving,” in Proceedings of the USENIX Annual Technical Conference (USENIX ATC ’21) , 2021
2021
Earlier work this paper cites.
T. Tambe, C. Hooper, L. Pentecost, T. Jia, E.-Y. Yang, M. Donato, V. Sanh, P. Whatmough, A. M. Rush, D. Brooks, and G.-Y. Wei, “EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference,” in Proceedings of the 54th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’21) , 2021
2021
Earlier work this paper cites.
C. Zhang, A. G. Kumbhare, I. Manousakis, D. Zhang, P. A. Misra, R. Assis, K. Woolcock, N. Mahalingam, B. Warrier, D. Gauthier, L. Kunnath, S. Solomon, O. Morales, M. Fontoura, and R. Bianchini, “Flex: High-Availability Datacenters with Zero Reserved Power,” in Proceedings of the 48th Annual International Symposium on Computer Architecture (ISCA ’21) , 2021
2021
Earlier work this paper cites.
Y. Zhang, W. Hua, Z. Zhou, G. E. Suh, and C. Delimitrou, “Sinan: ML-Based and QoS-Aware Resource Management for Cloud Microservices,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’21) , 2021
2021
Earlier work this paper cites.
S. Chen, A. Jin, C. Delimitrou, and J. F. Martínez, “ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud,” in Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA ’22) , 2022
2022
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” 2022
2022
Cited alongside, same era.
S. M. Nabavinejad, S. Reda, and M. Ebrahimi, “Coordinated batching and dvfs for dnn inference on gpu accelerators,” IEEE Transactions on Parallel and Distributed Systems , vol. 33, no. 10, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
F. Wang, W. Zhang, S. Lai, M. Hao, and Z. Wang, “Dynamic GPU Energy Optimization for Machine Learning Training Workloads,” IEEE Transactions on Parallel and Distributed Systems , 2022
2022
Cited alongside, same era.
M. Lammertyn, “60+ ChatGPT Statistics And Facts You Need to Know in 2024,” https://blog.invgate.com/chatgpt-statistics , 2024
2024
Closest in time.
LCG Consulting, “Energy Online: ERCOT Real Time Price,” https://energyonline.com/Data/GenericData.aspx?DataId=4&ERCOT___Real-time_Price , 2024
2024
Closest in time.
Meta, “Llama2-13B,” https://huggingface.co/meta-llama/Llama-2-13b , 2024
2024
Closest in time.
Meta, “Llama2-70B,” https://huggingface.co/meta-llama/Llama-2-70b , 2024
2024
Closest in time.
Meta, “Llama3-70B,” https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A Distributed Serving System for Transformer-Based Generative Models,” in Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’22) , 2022
2022
Cited alongside, same era.
Z. Zhou, X. Wei, J. Zhang, and G. Sun, “PetS: A unified framework for Parameter-Efficient transformers serving,” in Proceedings of the USENIX Annual Technical Conference (USENIX ATC ’22) , 2022
2022
Cited alongside, same era.
A. Agrawal, A. Panwar, J. Mohan, N. Kwatra, B. S. Gulavani, and R. Ramjee, “SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills,” 2023
2023
Cited alongside, same era.
A. Jahanshahi, M. Rezvani, and D. Wong, “WattWiser: Power and Resource-Efficient Scheduling for Multi-Model Multi-GPU Inference Servers,” in Proceedings of the 14th International Green and Sustainable Computing Conference (IGSC ’23) , 2023
2023
Cited alongside, same era.
Y. Jin, C.-F. Wu, D. Brooks, and G.-Y. Wei, “ S 3 S^{3} : Increasing GPU utilization during generative inference for higher throughput,” in Proceedings of the Advances in Neural Information Processing Systems (NeurIPS ’23) , 2023
2023
Cited alongside, same era.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23) , 2023
2023
Cited alongside, same era.
Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica, “AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving,” in Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’23) , 2023
2023
Cited alongside, same era.
K. K. W. Ng, H. M. Demoulin, and V. Liu, “Paella: Low-latency Model Serving with Software-defined GPU Scheduling,” in Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23) , 2023
2023
Cited alongside, same era.
X. Miao, C. Shi, J. Duan, X. Xi, D. Lin, B. Cui, and Z. Jia, “SpotServe: Serving Generative Large Language Models on Preemptible Instances,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’24) , 2024
2024
Closest in time.
Microsoft Azure, “ND H100 v5-series,” https://learn.microsoft.com/en-us/azure/virtual-machines/nd-h100-v5-series , 2024
2024
Closest in time.
Mistral AI, “The Mixtral-8x22B Large Language Model,” https://huggingface.co/mistralai/Mixtral-8x22B-Instruct-v0.1 , 2024
2024
Closest in time.
Mistral AI, “The Mixtral-8x7B Large Language Model,” https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1 , 2024
2024
Closest in time.
NVIDIA, “DGX H100: AI for Enterprise,” https://www.nvidia.com/en-us/data-center/dgx-h100/ , 2024
2024
Closest in time.
NVIDIA, “NVLink and NVLink Switch,” https://www.nvidia.com/en-us/data-center/nvlink/ , 2024
2024
Closest in time.
NVIDIA, “System Management Interface SMI,” https://developer.nvidia.com/system-management-interface , 2024
2024
Closest in time.
NVIDIA, “TensorRT-LLM’s Documentation,” https://nvidia.github.io/TensorRT-LLM/ , 2024
2024
Closest in time.
H. Oh, K. Kim, J. Kim, S. Kim, J. Lee, D.-s. Chang, and J. Seo, “ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’24) , 2024
2024
Closest in time.
P. Patel, E. Choukse, C. Zhang, I. Goiri, B. Warrier, N. Mahalingam, and R. Bianchini, “Characterizing Power Management Opportunities for LLMs in the Cloud,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’24) , 2024
2024
Closest in time.
P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative LLM inference using phase splitting,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24) , 2024
2024
Closest in time.
Python PuLP, “Optimization with PuLP,” https://coin-or.github.io/pulp/ , 2024
2024
Closest in time.
Python SciPy, “SciPy Library,” https://scipy.org/ , 2024
2024
Closest in time.
H. Qiu, W. Mao, A. Patke, S. Cui, S. Jha, C. Wang, H. Franke, Z. T. Kalbarczyk, T. Başar, and R. K. Iyer, “Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction,” in The 5th International Workshop on Cloud Intelligence / AIOps at ASPLOS 2024 , 2024
2024
Closest in time.
2024
Closest in time.
J. Stojkovic, N. Iliakopoulou, T. Xu, H. Franke, and J. Torrellas, “EcoFaaS: Rethinking the Design of Serverless Environments for Energy Efficiency,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24) , 2024
2024
Closest in time.
J. Stojkovic, P. Misra, I. Goiri, S. Whitlock, E. Choukse, M. Das, C. Bansal, J. Lee, Z. Sun, H. Qiu, R. Zimmermann, S. Samal, B. Warrier, A. Raniwala, and R. Bianchini, “SmartOClock: Workload- and Risk-Aware Overclocking in the Cloud,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24) , 2024
2024
Closest in time.
K. Talamadupula, “A Guide to LLM Inference Performance Monitoring,” https://symbl.ai/developers/blog/a-guide-to-llm-inference-performance-monitoring/ , 2024
2024
Closest in time.
Technology Innovation Institute (TII), “Falcon-180B,” 2024
2024
Closest in time.
T. Varshney, “Build an llm-powered data agent for data analysis,” https://developer.nvidia.com/blog/build-an-llm-powered-data-agent-for-data-analysis/ , 2024
2024
Closest in time.
Y. Zhang, Q. Wang, Z. Lin, P. Xu, and B. Wang, “Improving gpu energy efficiency through an application-transparent frequency scaling policy with performance assurance,” in Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys ’24) , 2024
2024
Closest in time.
Y. Z. Zhao, D. W. Wu, and J. Wang, “ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24) , 2024
2024
Closest in time.