Fetching the paper…
Reading the bibliography…
Model parallelism has become a necessity for training modern large-scale deep language models.
Performance analysis of a pipelined backpropagation parallel algorithm
Petrowski, A., Dreyfus, G., and Girault, C · 1993
Earlier work this paper cites.
Finding optimum wavefront of parallel computation
Sinharoy, B. and Szymanski, B · 1994
Earlier work this paper cites.
Scheduling of wavefront parallelism on scalable shared-memory multiprocessors
Manjikian, N. and Abdelrahman, T. S · 1996
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C · 2003
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A · 2010
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Optimizing performance of recurrent neural networks on gpus
Appleyard, J., Kocisky, T., and Blunsom, P · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Pipedream: Fast and efficient pipeline parallel dnn training
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., Devanur, N., Ganger, G., and Gibbons, P · 2018
Cited alongside, same era.
Exploring hidden dimensions in parallelizing convolutional neural networks
Jia, Z., Lin, S., Ruizhongtai Qi, C., and Aiken, A · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Later among the works it cites.
Supporting very large models using automatic dataflow graph partitioning
Wang, M., Huang, C.-c., and Li, J · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Dapple: A pipelined data parallel approach for training large models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Checkmate: Breaking the memory wall with optimal tensor rematerialization
Jain, P., Jain, A., Nrusimha, A., Gholami, A., Abbeel, P., Keutzer, K., Stoica, I., and Gonzalez, J. E · 2019
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Zero: Memory optimization towards training a trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I
Cited in the paper.
Fan, S., Rong, Y., Meng, C., Cao, Z., Wang, S., Zheng, Z., Wu, C., Long, G., Yang, J., Xia, L., et al · 2020
Later among the works it cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
The nvidia collective communication library (nccl)
NCCL · 2021
Closest in time.
Zero-offload: Democratizing billion-scale model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Closest in time.