Fetching the paper…
Reading the bibliography…
Due to the substantial computational cost, training state-of-the-art deep neural networks for large-scale datasets often requires distributed training using multiple computation workers.
Optimization of collective communication operations in MPICH
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
A simple, pipelined algorithm for large, irregular all-gather problems
J. L. Träff, A. Ripke, C. Siebert, P. Balaji, R. Thakur, and W. Gropp · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
J. Ba and D. Kingma · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
N. Strom · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Chainer: a next-generation open source framework for deep learning
S. Tokui, K. Oono, S. Hido, and J. Clayton · 2015
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
N. Dryden, S. A. Jacobs, T. Moon, and B. Van Essen · 2016
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Later among the works it cites.
Communication-efficient stochastic gradient descent, with applications to neural networks
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Later among the works it cites.
Automated inference with adaptive batches
S. De, A. Yadav, D. Jacobs, and T. Goldstein · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.