Fetching the paper…
Reading the bibliography…
The practical performance of online stochastic gradient descent algorithms is highly dependent on the chosen step size, which must be tediously hand-tuned in many applications.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
Harold J. Kushner and Dean S. Clark · 1978
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Cited alongside, same era.
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Later among the works it cites.
Gradient-based hypermarameter optimization through reversible learning
Douglas Maclaurin, David Duvenaud, and Ryan Adams · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…