Fetching the paper…
Reading the bibliography…
Stochastic gradient algorithms have been the main focus of large-scale learning problems and they led to important successes in machine learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker and Yann Le Cun · 1988
Earlier work this paper cites.
Automatic learning rate maximization by on-line estimation of the hessian’s eigenvectors
Yann LeCun, Patrice Y Simard, and Barak Pearlmutter · 1993
Earlier work this paper cites.
Directional newton methods in n variables
Yuri Levin and Adi Ben-Israel · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Directional secant method for nonlinear equations
Heng-Bin An and Zhong-Zhi Bai · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Theano: new features and speed improvements
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. Goodfellow, A. Bergeron, N. Bouchard, D. Warde-Farley, and Y. Bengio · 2012
Cited alongside, same era.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Cited alongside, same era.
Tom Schaul, Sixin Zhang, and Yann LeCun · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Later among the works it cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
Tom Schaul and Yann LeCun · 2013
Later among the works it cites.
Variance reduction for stochastic gradient optimization
Chong Wang, Xi Chen, Alex Smola, and Eric Xing · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. J. Goodfellow, D. Warde-Farley, P. Lamblin, V. Dumoulin, M. Mirza, R. Pascanu, J. Bergstra, F. Bastien, and Y. Bengio · 2013
Cited alongside, same era.