Fetching the paper…
Reading the bibliography…
When applied to training deep neural networks, stochastic gradient descent (SGD) often incurs steady progression phases, interrupted by catastrophic episodes in which loss and gradient norm explode.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Continuous inspection schemes
E. Page · 1954
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
Inference about the change point from cumulative sum-tests
D. Hinkley · 1970
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O(1/sqr(k))
Y. Nesterov · 1983
Earlier work this paper cites.
Increased rates of convergence through learning rate adaptation
R. A. Jacobs · 1988
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Le Cun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
N. Hansen and A. Ostermeier · 2001
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Adaptive stepsizes for recursive estimation with applications in approximate dynamic programming
A. P. George and W. B. Powell · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Cited alongside, same era.
Construction and Analysis of a Large Scale Image Ontology
J. Deng, K. Li, M. Do, H. Su, and L. Fei-Fei · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2010
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. Courville, P.-A. Manzagol, P. Vincent, and S. Bengio · 2010
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
No More Pesky Learning Rates
T. Schaul, S. Zhang, and Y. Le Cun · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. R. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Later among the works it cites.
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning via hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Cited alongside, same era.
Estimating the hessian by back-propagating curvature
J. Martens, I. Sutskever, and K. Swersky · 2012
Cited alongside, same era.
Lecture 6.5 - RMSProp, COURSERA: Neural Networks for Machine Learning
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Later among the works it cites.
Natural evolution strategies
D. Wierstra, T. Schaul, T. Glasmachers, Y. Sun, J. Peters, and J. Schmidhuber · 2014
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. G. Colmenarejo, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Y. Ollivier, L. Arnold, A. Auger, and N. Hansen · 2017
Closest in time.