Fetching the paper…
Reading the bibliography…
Training extremely large language models (LLMs) with billions of parameters is a computationally intensive task that pushes the limits of current data parallel training systems.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Sergeev, A. and Balso, M. D · 2018
Earlier work this paper cites.
Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect, 2019
Li, A., Song, S. L., Chen, J., Li, J., Liu, X., Tallent, N., and Barker, K · 2019
Earlier work this paper cites.
Performance analysis of deep learning workloads on leading-edge systems, 2019
Ren, Y., Yoo, S., and Hoisie, A · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
A unified architecture for accelerating distributed DNN training in heterogeneous GPU/CPU clusters
Jiang, Y., Zhu, Y., Lan, C., Yi, B., Cui, Y., and Guo, C · 2020
Earlier work this paper cites.
Pytorch distributed: experiences on accelerating data parallel training
Li, S., Zhao, Y., Varma, R., Salpekar, O., Noordhuis, P., Li, T., Paszke, A., Smith, J., Vaughan, B., Damania, P., and Chintala, S · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2020
Cited alongside, same era.
{ \{ Zero-offload } \} : Democratizing { \{ billion-scale } \} model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Cited alongside, same era.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
The falcon series of open language models, 2023
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Étienne Goffinet, Hesslow, D., Launay, J., Malartic, Q., Mazzotta, D., Noune, B., Pannier, B., and Penedo, G · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Zero++: Extremely efficient collective communication for giant model training
Wang, G., Qin, H., Jacobs, S. A., Holmes, C., Rajbhandari, S., Ruwase, O., Yan, F., Yang, L., and He, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mics: Near-linear scaling for training gigantic model on public cloud, 2022
Zhang, Z., Zheng, S., Wang, Y., Chiu, J., Karypis, G., Chilimbi, T., Li, M., and Jin, X · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Infiniband switching
NVIDIA
Cited in the paper.
Nvlink high-speed interconnect
NVIDIA
Cited in the paper.
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Later among the works it cites.
Communication-efficient large-scale distributed deep learning: A comprehensive survey, 2024
Liang, F., Zhang, Z., Lu, H., Leung, V. C. M., Guo, Y., and Hu, X · 2024
Closest in time.