Fetching the paper…
Reading the bibliography…
Recent innovation in large language models (LLMs), and their myriad use-cases have rapidly driven up the compute capacity demand for datacenter GPUs.
“Ensemble-level power management for dense blade servers”
Parthasarathy Ranganathan, Phil Leech, David Irwin and Jeffrey Chase · 2006
Earlier work this paper cites.
“Power Provisioning for a Warehouse-sized Computer”
Xiaobo Fan, Wolf-Dietrich Weber and Luiz Barroso · 2007
Earlier work this paper cites.
“Statistical profiling-based techniques for effective power provisioning in data centers”
Sriram Govindan, Jeonghwan Choi, Bhuvan Urgaonkar, Anand Sivasubramaniam and Andrea Baldini · 2009
Earlier work this paper cites.
“How much power oversubscription is safe and allowed in data centers”
Xing Fu, Xiaorui Wang and Charles Lefurgy · 2011
Earlier work this paper cites.
“Reverse engineering power management on NVIDIA GPUs - A detailed overview”
Martin Peres · 2013
Earlier work this paper cites.
“Power Capping: What Works, What Does Not”
Pavlos Petoumenos, Lev Mukhanov, Zheng Wang, Hugh Leather and Dimitrios. Nikolopoulos · 2015
Earlier work this paper cites.
“Dynamo: Facebook’s Data Center-Wide Power Management System”
Qiang Wu, Qingyuan Deng, Lakshmi Ganesh, Chang-Hong Hsu, Yun Jin, Sanjeev Kumar, Bin Li, Justin Meza and Yee Song · 2016
Earlier work this paper cites.
“Characterizing temperature, power, and soft-error behaviors in data center systems: Insights, challenges, and opportunities”
Bin Nie, Ji Xue, Saurabh Gupta, Christian Engelmann, Evgenia Smirni and Devesh Tiwari · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“SmoothOperator: Reducing Power Fragmentation and Improving Power Utilization in Large-Scale Datacenters”
Chang-Hong Hsu, Qingyuan Deng, Jason Mars and Lingjia Tang · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Towards Power Efficiency in Deep Learning on Data Center Hardware”
Miro Hodak, Masha Gorkovenko and Ajay Dholakia · 2019
Earlier work this paper cites.
“Analysis of Large-Scale Multi-Tenant GPU clusters for DNN training workloads”
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao and Fan Yang · 2019
Earlier work this paper cites.
“A Scalable Priority-aware Approach to Managing Data Center Server Power”
Yang Li, Charles Lefurgy, Karthick Rajamani, Malcolm Allen-Ware, Guillermo Silva, Daniel Heimsoth, Saugata Ghose and Onur Mutlu · 2019
Earlier work this paper cites.
“RoBERTa: A Robustly Optimized BERT Pretraining Approach”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer and Veselin Stoyanov · 2019
Earlier work this paper cites.
“Comparing GPU Power and Frequency Capping: A Case Study with the MuMMI Workflow”
Tapasya Patki, Zachary Frye, Harsh Bhatia, Francesco Di, James Glosli, Helgi Ingolfsson and Barry Rountree · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Cited alongside, same era.
“Why Your AI infrastructure Needs Both Training and Inference”, 2019
Tirias Research · 2019
Cited alongside, same era.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper and Bryan Catanzaro · 2019
Cited alongside, same era.
“GPU-NEST: Characterizing energy efficiency of multi-GPU inference servers”
Ali Jahanshahi, Hadi Sabzi, Chester Lau and Daniel Wong · 2020
Cited alongside, same era.
“Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at Scale”
Shaohong Li · 2020
Cited alongside, same era.
“Accelerate: Training and Inference at Scale Made Simple, Efficient and Adaptable.”, 2022
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller and Sourab Mangrulkar · 2022
Later among the works it cites.
“AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications”
Baolin Li, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Roy, Bill Bergeron, John Holodnak, Michael Houle, Matthew Hubbell, Michael Jones, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally and Devesh Tiwari · 2022
Later among the works it cites.
“Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model”
Alexandra Luccioni, Sylvain Viguier and Anne-Laure Ligozat · 2022
Later among the works it cites.
“Coordinated batching and DVFS for DNN inference on GPU accelerators”
Seyed Nabavinejad, Sherief Reda and Masoumeh Ebrahimi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan and Pritam Damania · 2020
Cited alongside, same era.
“Transformers: State-of-the-art Natural Language Processing”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest and Alexander. Rush · 2020
Cited alongside, same era.
“GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch”, 2021
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Shivanshu Purohit, Tri Songz, Wang Phil and Samuel Weinbach · 2021
Cited alongside, same era.
“On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”
Emily Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Cited alongside, same era.
“Characterization and prediction of deep learning workloads in large-scale gpu datacenters”
Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen and Tianwei Zhang · 2021
Cited alongside, same era.
“Prediction-Based Power Oversubscription in Cloud Platforms”
Alok Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde, Felipe Frujeri, Nithish Mahalingam, Pulkit Misra, Seyyed Javadi, Bianca Schroeder and Marcus Fontoura · 2021
Cited alongside, same era.
“Flex: High-Availability Datacenters With Zero Reserved Power”
Chaojie Zhang, Alok Kumbhare, Ioannis Manousakis, Deli Zhang, Pulkit Misra, Rod Assis, Kyle Woolcock, Nithish Mahalingam, Brijesh Warrier, David Gauthier, Lalu Kunnath, Steve Solomon, Osvaldo Morales, Marcus Fontoura and Ricardo Bianchini · 2021
Cited alongside, same era.
“The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink”
David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier and Jeff Dean · 2022
Later among the works it cites.
“BLOOM: A 176b-parameter open-access multilingual language model”
Teven Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Luccioni, François Yvon and Matthias Gallé · 2022
Later among the works it cites.
“Singularity: Planet-scale, preemptive and elastic scheduling of AI workloads”
Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha, Bhargav Gulavani, Rimma Nehme, Amey Agrawal, Chen Chen, Nipun Kwatra, Ramachandran Ramjee and Pankaj Sharma · 2022
Later among the works it cites.
“Not all GPUs are created equal: characterizing variability in large-scale, accelerator-rich systems”
Prasoon Sinha, Akhil Guliani, Rutwik Jain, Brandon Tran, Matthew Sinclair and Shivaram Venkataraman · 2022
Later among the works it cites.
Amazon Web Services
“Amazon SageMaker”, 2023 · 2023
Closest in time.
Microsoft Azure
“Azure Machine Learning - ML as a Service”, 2023 · 2023
Closest in time.
“EnvPipe: Performance-preserving DNN Training Framework for Saving Energy”
Sangjin Choi, Inhoe Koo, Jeongseob Ahn, Myeongjae Jeon and Youngjin Kwon · 2023
Closest in time.
“Large language models: fast proliferation and budding international competition”
International for Strategic · 2023
Closest in time.
“Towards Improved Power Management in Cloud GPUs”
Pratyush Patel, Zibo Gong, Syeda Rizvi, Esha Choukse, Pulkit Misra, Thomas Anderson and Akshitha Sriraman · 2023
Closest in time.
Google Cloud
“Vertex AI”, 2023 · 2023
Closest in time.
“Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training”
Jie You, Jae-Won Chung and Mosharaf Chowdhury · 2023
Closest in time.
“Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving”
Junyeol Yu, Jongseok Kim and Euiseong Seo · 2023
Closest in time.