Fetching the paper…
Reading the bibliography…
Recently, various distributed strategies for large language model training have been proposed.
A three-dimensional approach to parallel matrix multiplication
Agarwal, R. C., Balle, S. M., Gustavson, F. G., Joshi, M., and Palkar, P · 1995
Earlier work this paper cites.
Two-tree algorithms for full bandwidth broadcast, reduction and scan
Sanders, P., Speck, J., and Träff, J. L · 2009
Earlier work this paper cites.
Communication-optimal parallel 2.5d matrix multiplication and lu factorization algorithms
Solomonik, E. and Demmel, J · 2011
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Earlier work this paper cites.
Bringing hpc techniques to deep learning, 2017
Andrew, G · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes
Jia, X., Song, S., He, W., Wang, Y., Rong, H., Zhou, F., Xie, L., Guo, Z., Yang, Y., Yu, L., et al · 2018
Earlier work this paper cites.
Massively distributed sgd: Imagenet/resnet-50 training in a flash
Mikami, H., Suganuma, H., U-chupala, P., Tanaka, Y., and Kageyama, Y · 2018
Earlier work this paper cites.
Improving all-reduce collective operations for imbalanced process arrival patterns
Proficz, J · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow, 2018
Sergeev, A. and Balso, M. D · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2020
Cited alongside, same era.
Blink: Fast and generic collectives for distributed ml
Wang, G., Venkataraman, S., Phanishayee, A., Devanur, N., Thelin, J., and Stoica, I · 2020
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes, 2020
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Mics: Near-linear scaling for training gigantic model on public cloud
Zhang, Z., Zheng, S., Wang, Y., Chiu, J., Karypis, G., Chilimbi, T., Li, M., and Jin, X · 2022
Later among the works it cites.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2023
Closest in time.
Colossal-ai: A unified deep learning system for large-scale parallel training
Li, S., Liu, H., Bian, Z., Fang, J., Huang, H., Liu, Y., Wang, B., and You, Y · 2023
Closest in time.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maximizing parallelism in distributed training for huge neural networks, 2021
Bian, Z., Xu, Q., Wang, B., and You, Y · 2021
Cited alongside, same era.
1-bit lamb: Communication efficient large-scale large-batch training with lamb’s convergence speed, 2021
Li, C., Awan, A. A., Tang, H., Rajbhandari, S., and He, Y · 2021
Cited alongside, same era.
2.5-dimensional distributed model training
Wang, B., Xu, Q., Bian, Z., and You, Y · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Zero++: Extremely efficient collective communication for giant model training, 2023
Wang, G., Qin, H., Jacobs, S. A., Holmes, C., Rajbhandari, S., Ruwase, O., Yan, F., Yang, L., and He, Y · 2023
Closest in time.
G-meta: Distributed meta learning in gpu clusters for large-scale recommender systems
Xiao, Y., Zhao, S., Zhou, Z., Huan, Z., Ju, L., Zhang, X., Wang, L., and Zhou, J · 2023
Closest in time.
An efficient 2d method for training super-large deep learning models
Xu, Q. and You, Y · 2023
Closest in time.
Pytorch fsdp: Experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Closest in time.