J. Kaplan et al. , “Scaling Laws for Neural Language Models,” 2020, _eprint: 2001.08361
2001
Earlier work this paper cites.
O. Sharir, B. Peleg, and Y. Shoham, “The Cost of Training NLP Models: A Concise Overview,” 2020, _eprint: 2004.08900
2004
Earlier work this paper cites.
T.B. Brown et al. , “Language Models are Few-Shot Learners,” 2020, _eprint: 2005.14165
2005
Earlier work this paper cites.
N.C. Thompson et al. , “The Computational Limits of Deep Learning,” 2020, _eprint: 2007.05558
2007
Earlier work this paper cites.
D. Amodei and D. Hernandez, “ AI and Compute ,” May 2018, published: OpenAI Blog
2018
Earlier work this paper cites.
M. Naumov et al. , “Deep Learning Recommendation Model for Personalization and Recommendation Systems,” May 2019
2019
Earlier work this paper cites.
M. Shoeybi et al. , “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism,” Sep. 2019
2019
Earlier work this paper cites.
Z. Lan et al. , “ALBERT: A Lite BERT for Self-supervised Learning of Language Representations,” Sep. 2019
2019
Earlier work this paper cites.
D. Lepikhin et al. , “Gshard: Scaling giant models with conditional computation and automatic sharding,” 2020
2020
Earlier work this paper cites.
N. Ahmed and M. Wahed, “The de-democratization of ai: Deep learning and the compute divide in artificial intelligence research,” 2020
2020
Earlier work this paper cites.
C. Rosset, “Turing-NLG: A 17-billion-parameter language model by Microsoft,” Feb. 2020
2020
Earlier work this paper cites.
Z. Zhang et al. , “CPM: A Large-scale Generative Chinese Pre-trained Language Model,” Dec. 2020
2020
Earlier work this paper cites.
W. Antoun, F. Baly, and H. Hajj, “AraGPT2: Pre-Trained Transformer for Arabic Language Generation,” Dec. 2020
2020
Earlier work this paper cites.