Fetching the paper…
Reading the bibliography…
Recent work has established an empirically successful framework for adapting learning rates for stochastic gradient descent (SGD).
Increased rates of convergence through learning rate adaptation
Jacobs, R. A · 1988
Earlier work this paper cites.
Efficient backprop
LeCun, Y, Bottou, L, Orr, G, and Muller, K · 1998
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
Almeida, L and Langlois, T · 1999
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Amari, S, Park, H, and Fukumizu, K · 2000
Earlier work this paper cites.
Adaptive stepsizes for recursive estimation with applications in approximate dynamic programming
George, A. P and Powell, W. B · 2006
Cited alongside, same era.
Topmoumoute online natural gradient algorithm, 2008
Le Roux, N, Manzagol, P, and Bengio, Y · 2008
Cited alongside, same era.
Sgd-qn: Careful quasi-newton stochastic gradient descent
Bordes, A, Bottou, L, and Gallinari, P · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C, Hazan, E, and Singer, Y · 2010
Cited alongside, same era.
A fast natural Newton method
Nicolas Le Roux, A. F
Cited in the paper.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Niu, F, Recht, B, Re, C, and Wright, S. J · 2011
Later among the works it cites.
No More Pesky Learning Rates
Schaul, T, Zhang, S, and LeCun, Y · 2012
Later among the works it cites.
Sample size selection in optimization methods for machine learning
Byrd, R, Chin, G, Nocedal, J, and Wu, Y · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…