PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Original
Harlap, A., Narayanan, D., Phanishayee, A., Seshadri, V., Devanur, N., Ganger, G., and Gibbons, P · 2018
Cited alongside, same era.
Gist: Efficient Data Encoding for Deep Neural Network Training
Jain, A., Phanishayee, A., Mars, J., Tang, L., and Pekhimenko, G · 2018
Cited alongside, same era.
Beyond Data and Model Parallelism for Deep Neural Networks
Jia, Z., Zaharia, M., and Aiken, A · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Cited alongside, same era.
PipeDream: Generalized Pipeline Parallelism for DNN Training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
https://github.com/nvidia/megatron-lm
Megatron Repository
Cited in the paper.
https://github.com/NVIDIA/DeepLearningExamples/blob/master/PyTorch/LanguageModeling/BERT/README.md#results
NVIDIA Deep Learning Examples, BERT
Cited in the paper.
https://github.com/jcpeterson/openwebtext
OpenWebText Dataset
Cited in the paper.