Fetching the paper…
Reading the bibliography…
Synchronized stochastic gradient descent (SGD) optimizers with data parallelism are widely used in training large-scale deep neural networks.
A Simple Weight Decay Can Improve Generalization
Anders Krogh and John A. Hertz. 1992 · 1992
Earlier work this paper cites.
Interprocessor collective communication library (InterCom). In Proceedings of IEEE Scalable High Performance Computing Conference
M. Barnett, L. Shuler, R. van de Geijn, S. Gupta, D. G. Payne, and J. Watts. 1994 · 1994
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Y Bengio, and Jean-Pierre David. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
Jianmin Chen Zhifeng Chen Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Geoffrey Irving Michael Isard Manjunath Kudlur Josh Levenberg Rajat Monga Sherry Moore Derek G. Murray Benoit Steiner Paul Tucker Vijay Vasudevan Pete Warden Martin Wicke Yuan Yu Xiaoqiang Zheng Martin Abadi, Paul Barham. 2015 · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning.. In OSDI
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, and others. 2016 · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Cited alongside, same era.
Ako: Decentralised deep learning with partial gradient exchange. In Proceedings of the Seventh ACM Symposium on Cloud Computing
Pijika Watcharapichat, Victoria Lopez Morales, Raul Castro Fernandez, and Peter Pietzuch. 2016 · 2016
Cited alongside, same era.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda. 2017 · 2017
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Later among the works it cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaev, Ganesh Venkatesh, and others. 2017 · 2017
Later among the works it cites.
Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs
Shaohuai Shi and Xiaowen Chu. 2017 · 2017
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minsik Cho, Ulrich Finkler, Sameer Kumar, David Kung, Vaibhav Saxena, and Dheeraj Sreedhar. 2017 · 2017
Cited alongside, same era.
Valeriu Codreanu, Damian Podareanu, and Vikram Saletore. 2017 · 2017
Cited alongside, same era.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland. 2017 · 2017
Cited alongside, same era.
Bringing HPC techniques to deep learning
Andrew Gibiansky. 2017 · 2017
Cited alongside, same era.
Yang You, Igor Gitman, and Boris Ginsburg. 2017a · 2017
Later among the works it cites.
Yang You, Zhao Zhang, C Hsieh, James Demmel, and Kurt Keutzer. 2017b · 2017
Later among the works it cites.
High-Accuracy Low-Precision Training
Christopher De Sa, Megan Leszczynski, Jian Zhang, Alana Marzoev, Christopher R Aberger, Kunle Olukotun, and Christopher Ré. 2018 · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Closest in time.
Shaohuai Shi, Qiang Wang, Xiaowen Chu, and Bo Li. 2018 · 2018
Closest in time.