Fetching the paper…
Reading the bibliography…
Stochastic gradient methods (SGMs) are the predominant approaches to train deep learning models.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some aspects of parallel and distributed iterative algorithms—a survey
D. P. Bertsekas and J. N. Tsitsiklis · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Libsvm: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. Le, and A. Ng · 2012
Earlier work this paper cites.
RMSProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
An asynchronous parallel stochastic coordinate descent algorithm
J. Liu, S. Wright, C. Ré, V. Bittorf, and S. Sridhar · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Cited alongside, same era.
An asynchronous mini-batch algorithm for regularized stochastic optimization
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
Revisiting small batch training for deep neural networks
D. Masters and C. Luschi · 2018
Later among the works it cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
Error compensated quantized sgd and its applications to large-scale distributed optimization
J. Wu, W. Huang, J. Huang, and T. Zhang · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Y. Yan, T. Yang, Z. Li, Q. Lin, and Y. Yang · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arock: an algorithmic framework for asynchronous parallel coordinate updates
Z. Peng, Y. Xu, M. Yan, and W. Yin · 2016
Cited alongside, same era.
Adadelay: Delay adaptive distributed stochastic optimization
S. Sra, A. W. Yu, M. Li, and A. Smola · 2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
A downsampled variant of imagenet as an alternative to the cifar datasets
P. Chrabaszcz, I. Loshchilov, and F. Hutter · 2017
Cited alongside, same era.
Delay compensated asynchronous adam algorithm for deep neural networks
N. Guan, L. Shan, C. Yang, W. Xu, and M. Zhang · 2017
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
H. Mania, X. Pan, D. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan · 2017
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
J. Chen and Q. Gu · 2018
Cited alongside, same era.
K. Bäckström, M. Papatriantafilou, and P. Tsigas · 2019
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2019
Later among the works it cites.
Convergence analyses of online adam algorithm in convex setting and two-layer relu neural network
B. Fang and D. Klabjan · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, and Y. Liu · 2019
Later among the works it cites.
Dadam: A consensus-based distributed adaptive gradient method for online optimization
P. Nazari, D. A. Tarzanagh, and G. Michailidis · 2019
Later among the works it cites.
On the convergence of asynchronous parallel iteration with unbounded delays
Z. Peng, Y. Xu, M. Yan, and W. Yin · 2019
Later among the works it cites.
On the convergence proof of amsgrad and a new version
P. T. Tran et al · 2019
Later among the works it cites.
Asynchronous gradient push
M. S. Assran and M. G. Rabbat · 2020
Closest in time.
Distributed learning systems with first-order methods
J. Liu, C. Zhang, et al · 2020
Closest in time.
SAdam: A variant of adam for strongly convex functions
G. Wang, S. Lu, Q. Cheng, W. Tu, and L. Zhang · 2020
Closest in time.