Fetching the paper…
Reading the bibliography…
The costs of training frontier AI models have grown dramatically in recent years, but there is limited public data on the magnitude and growth of these expenses.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
The Datacenter As a Computer: Designing Warehouse-scale Machines, 2018
Luiz Andre Barroso, Urs Holzle, Parthasarathy Ranganathan, and Margaret Martonosi · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Ten Lessons From Three Generations Shaped Google’s TPUv4i : Industrial Product
Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson · 2021
Earlier work this paper cites.
Carbon Emissions and Large Neural Network Training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Earlier work this paper cites.
The importance of (exponentially more) computing power
Neil C Thompson, Shuning Ge, and Gabriel F Manso · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Compute trends across three eras of machine learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos · 2022
Earlier work this paper cites.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al · 2022
Earlier work this paper cites.
OPT: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Earlier work this paper cites.
Sustainable AI: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Earlier work this paper cites.
Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat · 2022
Earlier work this paper cites.
Artificial Intelligence Index Report 2023
Nestor Maslej, Loredana Fattorini, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Helen Ngo, Juan Carlos Niebles, Vanessa Parli, et al · 2023
Cited alongside, same era.
NVIDIA DGX SuperPOD Reference Architecture
NVIDIA Corporation · 2023
Cited alongside, same era.
NVIDIA DGX H100 Datasheet
NVIDIA Corporation · 2023
Cited alongside, same era.
NVIDIA DGX SuperPOD Data Center Design (for NVIDIA DGX H100 Systems)
NVIDIA Corporation · 2023
Cited alongside, same era.
What if Dario Amodei Is Right about A.I.?
Ezra Klein and Dario Amodei · 2024
Cited alongside, same era.
Artificial Intelligence Index Report 2024, 2024
Nestor Maslej, Loredana Fattorini, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Helen Ngo, Juan Carlos Niebles, Vanessa Parli, et al · 2024
Google discussed dropping Broadcom as their AI chips supplier, 2023
The Information · 2024
Closest in time.
AI server cost analysis – memory is the biggest loser, 2023
Dylan Patel and Gerald Wong · 2024
Closest in time.
Raymond James estimates it costs Nvidia $3,320
Tae Kim · 2024
Closest in time.
NVIDIA Tesla K80 Specs
TechPowerUp · 2024
Closest in time.
NVIDIA Tesla P100 PCIe Specs
TechPowerUp · 2024
Closest in time.
NVIDIA Tesla V100 SXM2 32 GB Specs
TechPowerUp · 2024
Closest in time.
NVIDIA A100 SXM4 40GB Specs
TechPowerUp · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parameter, Compute and Data Trends in Machine Learning
Epoch AI · 2024
Cited alongside, same era.
Trends in Machine Learning Hardware
Marius Hobbhahn, Lennart Heim, and Gökçe Aydos · 2024
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team · 2024
Cited alongside, same era.
How Much Does an Employee Cost You?
Barbara Weltman · 2024
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
The length of time spent training notable models is growing
Epoch AI · 2024
Cited alongside, same era.
Data centers - Meta sustainability
Meta · 2024
Closest in time.
AI Datacenter Energy Dilemma - Race for AI Datacenter Space
Dylan Patel, Daniel Nishball, and Jeremie Eliahou Ontiveros · 2024
Closest in time.
https://www.eia.gov/electricity/monthly/epm_table_grapher.php?t=epmt_5_6_a , 2024
Electric Power Monthly · 2024
Closest in time.
Why GPUs are great for AI
NVIDIA Corporation · 2024
Closest in time.
https://web.archive.org/web/20240407085026/https://www.eia.gov/energyexplained/electricity/electricity-in-the-us-top-10.php , 2022
Electricity generation, capacity, and sales in the United States · 2024
Closest in time.