Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exploded in popularity due to their new generative capabilities that go far beyond prior state-of-the-art.
J. McDonald, B. Li, N. Frey et al. , “Great power, great responsibility: Recommendations for reducing energy for training language models,” in Findings of the Association for Computational Linguistics: NAACL 2022 , 2022, pp. 1962–1970
1970
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar et al. , “Attention is all you need,” 2017
2017
Earlier work this paper cites.
A. Reuther, J. Kepner, C. Byun et al. , “Interactive supercomputing on 40,000 cores for machine learning and data analysis,” in 2018 IEEE High Performance extreme Computing Conference (HPEC) . IEEE, 2018, pp. 1–6
2018
Earlier work this paper cites.
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in NLP,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 3645–3650
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee et al. , “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://aclanthology.org/N19-1423
2019
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri et al. , “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020
2020
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” 2021
2021
Earlier work this paper cites.
FairScale authors, “Fairscale: A general purpose modular pytorch library for high performance and large scale training,” https://github.com/facebookresearch/fairscale
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Foster, Generative deep learning . ” O’Reilly Media, Inc.”, 2022
2022
Cited alongside, same era.
J. Sevilla, L. Heim, A. Ho et al. , “Compute trends across three eras of machine learning,” 2022
2022
Cited alongside, same era.
D. Zhao, N. C. Frey, J. McDonald et al. , “A green(er) world for a.i.” in 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) , 2022, pp. 742–750
2022
Cited alongside, same era.
B. Li, T. Patel, S. Samsi et al. , “Miso: exploiting multi-instance gpu capability on multi-tenant gpu clusters,” in Proceedings of the 13th Symposium on Cloud Computing , 2022, pp. 173–189
2022
Cited alongside, same era.
Stability-AI, “Stable Diffusion,” https://github.com/Stability-AI/StableDiffusion
2023
Cited alongside, same era.
D. Patel, “The ai brick wall – a practical limit for scaling dense transformer models, and how gpt 4 will break past it,” https://www.semianalysis.com/p/the-ai-brick-wall-a-practical-limit
2023
Closest in time.
R. Desislavov, F. Martínez-Plumed, and J. Hernández-Orallo, “Trends in ai inference energy consumption: Beyond the performance-vs-parameter laws of deep learning,” Sustainable Computing: Informatics and Systems , vol. 38, p. 100857, 2023
2023
Closest in time.
H. Touvron, T. Lavril, G. Izacard et al. , “Llama: Open and efficient foundation language models,” 2023
2023
Closest in time.
R. Gozalo-Brizuela and E. C. Garrido-Merchan, “Chatgpt is not all you need. a state of the art review of large generative ai models,” 2023
2023
Closest in time.
Facebook Research, online, 2023. [Online]. Available: {https://github.com/facebookresearch/llama}
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Epstein, A. Hertzmann, the Investigators of Human Creativity et al. , “Art and the science of generative ai,” Science , vol. 380, no. 6650, pp. 1110–1111, 2023. [Online]. Available: https://www.science.org/doi/abs/10.1126/science.adh4451
2023
Cited alongside, same era.
B. Meskó and E. J. Topol, “The imperative for regulatory oversight of large language models (or generative ai) in healthcare,” npj Digital Medicine , vol. 6, no. 1, p. 120, 2023
2023
Cited alongside, same era.
H. Zohny, J. McMillan, and M. King, “Ethics of generative ai,” Journal of Medical Ethics , vol. 49, no. 2, pp. 79–80, 2023. [Online]. Available: https://jme.bmj.com/content/49/2/79
2023
Cited alongside, same era.
“Nvidia/megatron-lm: Ongoing research training transformer models at scale,” https://github.com/NVIDIA/Megatron-LM
2023
Cited alongside, same era.
“Different development paths of llms,” https://www.interconnects.ai/p/llm-development-paths
Cited in the paper.
NVIDIA. Nvidia-smi. [Online]. Available: http://developer.download.nvidia.com/compute/DCGM/docs/nvidia-smi-367.38.pdf
Cited in the paper.
——. Nvidia data center GPU manager (dcgm). [Online]. Available: https://developer.nvidia.com/dcgm
Cited in the paper.
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang et al. , “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca
2023
Closest in time.
NVIDIA, “Multi-Process Service,” https://docs.nvidia.com/deploy/mps/
2023
Closest in time.
NVIDIA, “NVIDIA Multi Instance GPU User Guide,” https://docs.nvidia.com/datacenter/tesla/mig-user-guide/
2023
Closest in time.