Fetching the paper…
Reading the bibliography…
The rapid adoption of large language models (LLMs) has led to significant advances in natural language processing and text generation.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
A multiplier adjustment method for the generalized assignment problem
Marshall L Fisher, Ramchandran Jaikumar, and Luk N Van Wassenhove. 1986 · 1986
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Samuel Williams, Andrew Waterman, and David Patterson. 2009 · 2009
Earlier work this paper cites.
An energy efficient task scheduling scheme for heterogeneous GPU-enhanced clusters. In 2012 International Conference on Systems and Informatics (ICSAI2012) . IEEE, 623–627
Hongpeng Huo, Chongchong Sheng, Xinming Hu, and Baifeng Wu. 2012 · 2012
Earlier work this paper cites.
Modeling the energy efficiency of heterogeneous clusters. In 2014 43rd International Conference on Parallel Processing . IEEE, 321–330
Lavanya Ramapantulu, Bogdan Marius Tudor, Dumitrel Loghin, Trang Vu, and Yong Meng Teo. 2014 · 2014
Earlier work this paper cites.
A comparative study of resource allocation strategies for a green cloud. In 2016 2nd International Conference on Next Generation Computing Technologies (NGCT) . 621–625
Satveer and Mahendra Singh Aswal. 2016 · 2016
Earlier work this paper cites.
Energy efficient real-time task scheduling on CPU-GPU hybrid clusters. In IEEE INFOCOM 2017-IEEE Conference on Computer Communications . IEEE, 1–9
Xinxin Mei, Xiaowen Chu, Hai Liu, Yiu-Wing Leung, and Zongpeng Li. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, et al · 2017
Earlier work this paper cites.
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau. 2020 · 2020
Earlier work this paper cites.
CPU–GPU utilization aware energy-efficient scheduling algorithm on heterogeneous computing systems
Xiaoyong Tang and Zhuojun Fu. 2020 · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Characterization and prediction of deep learning workloads in large-scale GPU datacenters. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’21) . Association for Computing Machinery, New York, NY, USA, Article 104, 15 pages
Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen, and Tianwei Zhang. 2021 · 2021
Earlier work this paper cites.
Carbon Emissions and Large Neural Network Training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021 · 2021
Earlier work this paper cites.
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, and et al. 2022 · 2022
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, et al · 2022
Earlier work this paper cites.
Carbon-aware computing for datacenters
Ana Radovanović, Ross Koningstein, Ian Schneider, Bokan Chen, Alexandre Duarte, Binz Roy, Diyue Xiao, Maya Haridasan, Patrick Hung, Nick Care, et al · 2022
Cited alongside, same era.
DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.). PMLR, 18332–18346
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. 2022 · 2022
Cited alongside, same era.
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, and et al. 2022 · 2022
Cited alongside, same era.
Treehouse: A Case For Carbon-Aware Datacenter Software
Thomas Anderson, Adam Belay, Mosharaf Chowdhury, Asaf Cidon, and Irene Zhang. 2023 · 2023
Cited alongside, same era.
Open LLM Leaderboard
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Later among the works it cites.
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference. In 2023 IEEE High Performance Extreme Computing Conference (HPEC) . 1–9
Siddharth Samsi, Dan Zhao, Joseph McDonald, Baolin Li, Adam Michaleas, Michael Jones, William Bergeron, Jeremy Kepner, Devesh Tiwari, and Vijay Gadepally. 2023 · 2023
Later among the works it cites.
Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline. In Thirty-seventh Conference on Neural Information Processing Systems
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. 2023 · 2023
Cited alongside, same era.
Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning
Radosvet Desislavov, Fernando Martínez-Plumed, and José Hernández-Orallo. 2023 · 2023
Cited alongside, same era.
SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy Saving. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–13
Kaijie Fan, Marco D’Antonio, Lorenzo Carpentieri, Biagio Cosenza, Federico Ficarelli, and Daniele Cesarini. 2023 · 2023
Cited alongside, same era.
Energy-Efficient GPU Clusters Scheduling for Deep Learning
Diandian Gu, Xintong Xie, Gang Huang, Xin Jin, and Xuanzhe Liu. 2023 · 2023
Cited alongside, same era.
Towards Application Centric Carbon Emission Management. In Proceedings of the 2nd Workshop on Sustainable Computer Systems (Boston, MA, USA) (HotCarbon ’23) . Association for Computing Machinery, New York, NY, USA, Article 5, 7 pages
Sudarsun Kannan and Ulrich Kremer. 2023 · 2023
Cited alongside, same era.
Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference Service. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’23) . Association for Computing Machinery, New York, NY, USA, Article 20, 15 pages
Baolin Li, Siddharth Samsi, Vijay Gadepally, and Devesh Tiwari. 2023 · 2023
Cited alongside, same era.
Model-driven cluster resource management for ai workloads in edge clouds
Qianlin Liang, Walid A Hanafy, Ahmed Ali-Eldin, and Prashant Shenoy. 2023 · 2023
Cited alongside, same era.
Adapting Datacenter Capacity for Greener Datacenters and Grid. In Proceedings of the 14th ACM International Conference on Future Energy Systems (Orlando, FL, USA) (e-Energy ’23) . Association for Computing Machinery, New York, NY, USA, 200–213
Liuzixuan Lin and Andrew A Chien. 2023 · 2023
Cited alongside, same era.
Closest in time.
Exploding AI Power Use: an Opportunity to Rethink Grid Planning and Management. In Proceedings of the 15th ACM International Conference on Future and Sustainable Energy Systems (Singapore, Singapore) (e-Energy ’24) . Association for Computing Machinery, New York, NY, USA, 434–441
Liuzixuan Lin, Rajini Wijayawardana, Varsha Rao, Hai Nguyen, Emmanuel Wedan GNIBGA, and Andrew A. Chien. 2024 · 2024
Closest in time.
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Timothy R. McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N. Halgamuge. 2024 · 2024
Closest in time.
NVIDIA-NVML
NVIDIA. Accessed 2024 · 2024
Closest in time.
PyJoules: Python-based energy measurement library for various domains including NVIDIA GPUs
PowerAPI. 2024 · 2024
Closest in time.
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Jovan Stojkovic, Esha Choukse, Chaojie Zhang, Inigo Goiri, and Josep Torrellas. 2024 · 2024
Closest in time.
Stanford alpaca: An instruction following llama model
R. Taori, I. Gulrajani, T. Zhang, and et al. 2024 · 2024
Closest in time.
Gemini: A Family of Highly Capable Multimodal Models
Google Gemini Team. 2024 · 2024
Closest in time.
Towards Efficient and Reliable LLM Serving: A Real-World Workload Study
Yuxin Wang, Yuhan Chen, Zeyu Li, Zhenheng Tang, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, and Xiaowen Chu. 2024 · 2024
Closest in time.
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads. In Proceedings of the 15th ACM International Conference on Future and Sustainable Energy Systems (e-Energy ’24) . Association for Computing Machinery, New York, NY, USA, 506–513
Grant Wilkins, Srinivasan Keshav, and Richard Mortier. 2024 · 2024
Closest in time.
Reducing the Carbon Impact of Generative AI Inference (Today and in 2035). In Proceedings of the 2nd Workshop on Sustainable Computer Systems (Boston, MA, USA) (HotCarbon ’23) . Association for Computing Machinery, New York, NY, USA, Article 11, 7 pages
Andrew A Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana. 2023 · 2035
Closest in time.