Fetching the paper…
Reading the bibliography…
Optimization by stochastic gradient descent is an important component of many large-scale machine learning algorithms.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
A. Benveniste, M. Metivier, and P. Priouret · 1990
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S Sutton · 1992
Earlier work this paper cites.
Temporal-difference methods and Markov models
Etienne Barnard · 1993
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming
Yurii Nesterov and Arkadii Semenovich Nemirovskii · 1994
Earlier work this paper cites.
Online Algorithms and Stochastic Approximations
Léon Bottou · 1998
Earlier work this paper cites.
Efficient BackProp
Y. LeCun, L. Bottou, G. Orr, and K. Muller · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
The MNIST dataset of handwritten digits
Yann LeCun and Corinna Cortes · 1998
Earlier work this paper cites.
Large Scale Online Learning
Léon Bottou and Yann LeCun · 2004
Cited alongside, same era.
On the role of tracking in stationary environments
Richard S. Sutton, Anna Koop, and David Silver · 2007
Cited alongside, same era.
The Tradeoffs of Large Scale Learning
Léon Bottou and Olivier Bousquet · 2008
Cited alongside, same era.
Topmoumoute online natural gradient algorithm, 2008
N. Le Roux, P.A. Manzagol, and Y. Bengio · 2008
Cited alongside, same era.
SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Cited alongside, same era.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2010
Cited alongside, same era.
Towards Optimal One Pass Large Scale Learning with Averaged Stochastic Gradient Descent
Wei Xu · 2011
Later among the works it cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton · 2012
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Later among the works it cites.
ADADELTA: An Adaptive Learning Rate Method
Matthew D Zeiler · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Real-parameter black-box optimization benchmarking 2010: Experimental setup
Nikolaus Hansen, Anne Auger, Steffen Finck, Raymond Ros, et al · 2010
Cited alongside, same era.
Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009
Nikolaus Hansen, Anne Auger, Raymond Ros, Steffen Finck, and Petr Pošík · 2010
Cited alongside, same era.
Later among the works it cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T Tieleman and G Hinton · 2012
Later among the works it cites.
No More Pesky Learning Rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Closest in time.
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Closest in time.
Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
Tom Schaul and Yann LeCun · 2013
Closest in time.