Fetching the paper…
Reading the bibliography…
Training large neural networks requires distributing learning across multiple workers, where the cost of communicating gradients can be a significant bottleneck.
Sui confini della probabilità
Cantelli, F. P · 1928
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Note on a Method for Calculating Corrected Sums of Squares and Products
Welford, B. P · 1962
Earlier work this paper cites.
A Direct Adaptive Method for Faster Backpropagation Learning: the RPROP Algorithm
Riedmiller, M. and Braun, H · 1993
Earlier work this paper cites.
The Three Sigma Rule
Pukelsheim, F · 1994
Earlier work this paper cites.
The Art of Computer Programming, Volume 2 (3rd Ed.): Seminumerical Algorithms
Knuth, D. E · 1997
Earlier work this paper cites.
Information Theory, Inference & Learning Algorithms
MacKay, D. J. C · 2002
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Cubic Regularization of Newton Method and its Global Performance
Nesterov, Y. and Polyak, B · 2006
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A · 2009
Earlier work this paper cites.
Understanding the Difficulty of Training Deep Feedforward Neural Networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2013
Cited alongside, same era.
Identifying and Attacking the Saddle Point Problem in High-Dimensional Non-Convex Optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Scaling Distributed Machine Learning with the Parameter Server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Cited alongside, same era.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
Richtárik, P. and Takáč, M · 2014
Cited alongside, same era.
1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Cited alongside, same era.
Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
Akiba, T., Suzuki, S., and Fukuda, K · 2017
Later among the works it cites.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Later among the works it cites.
Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients
Balles, L. and Hennig, P · 2017
Later among the works it cites.
Why Momentum Really Works
Goh, G · 2017
Later among the works it cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
Deep Learning in Neural Networks: an Overview
Schmidhuber, J · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Strom, N · 2015
Cited alongside, same era.
Stochastic spectral descent for discrete graphical models
Carlson, D., Hsieh, Y., Collins, E., Carin, L., and Cevher, V · 2016
Cited alongside, same era.
Accelerated Gradient Descent Escapes Saddle Points Faster than Gradient Descent
Jin, C., Netrapalli, P., and Jordan, M. I · 2017
Later among the works it cites.
Fixing Weight Decay Regularization in Adam
Loshchilov, I. and Hutter, F · 2017
Later among the works it cites.
Distributed Mean Estimation with Limited Communication
Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, H. B · 2017
Later among the works it cites.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Later among the works it cites.
On the Convergence of Adam and Beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Closest in time.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Smith, S. L. and Le, Q. V · 2018
Closest in time.