Fetching the paper…
Reading the bibliography…
We propose a stochastic modified equations (SME) for modeling the asynchronous stochastic gradient descent (ASGD) algorithms.
Sur les fonctions absolument monotones
Serge Bernstein · 1929
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Stochastic differential equations
P.E. Kloeden and E. Platen · 1992
Earlier work this paper cites.
Verification and Planning for Stochastic Processes with Asynchronous Events
HL Younes · 2005
Earlier work this paper cites.
Distributed asynchronous online learning for natural language processing
K. Gimpel, D. Das, and N.A. Smith · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J.C. Duchi · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J.C. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Ré, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Estimation, optimization, and parallelism when data is sparse
J.C. Duchi, M.I. Jordan, and B. McMahan · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
D. Needell, R. Ward, and N. Srebro · 2014
Cited alongside, same era.
Stochastic processes and applications: Diffusion processes, the Fokker-Planck and Langevin equations
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
J. Liu and S.J. Wright · 2015
Later among the works it cites.
An asynchronous parallel stochastic coordinate descent algorithm
J. Liu, S.J. Wright, C. Ré, V. Bittorf, and S. Sridhar · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning, 2016
L. Bottou, F.E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
Asynchrony begets momentum, with an application to deep learning
I. Mitliagkas, C. Zhang, S. Hadjis, and C. Ré · 2016
Later among the works it cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G.A. Pavliotis · 2014
Cited alongside, same era.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
P. Richtárik and M. Takáč · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D.P. Kingma and J. Ba · 2015
Cited alongside, same era.
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N.S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and W. E · 2017
Later among the works it cites.