Fetching the paper…
Reading the bibliography…
We propose a statistical adaptive procedure called SALSA for automatically scheduling the learning rate (step size) in stochastic gradient methods.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Accelerated stochastic approximation
Harry Kesten · 1958
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
A stochastic analog of the conjugate gradient method
A. M. Gupal and L. T. Bazhenov · 1972
Earlier work this paper cites.
Comparison of the rates of convergence of one-step and multi-step optimization algorithms in the presence of noise
Boris T. Polyak · 1977
Earlier work this paper cites.
Adaptive step adjustment for a stochastic optimization algorithm
F. Mirzoakhmedov and S. P. Uryasev · 1983
Earlier work this paper cites.
On the determination of the step size in stochastic quasigradient methods
Georg Ch. Pflug · 1983
Earlier work this paper cites.
Stochastic approximation method with gradient averaging for unconstrained problems
Andrzej Ruszczyński and Wojciech Syski · 1983
Earlier work this paper cites.
Stochastic approximation algorithm with gradient averaging and on-line stepsize rules
Andrzej Ruszczyński and Wojciech Syski · 1984
Earlier work this paper cites.
Increased rates of convergence through learning rate adaption
R. A. Jacobs · 1988
Earlier work this paper cites.
Adaptive stepsize control in stochastic approximation algorithms
Georg Ch. Pflug · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of Delta-Bar-Delta
Richard S. Sutton · 1992
Earlier work this paper cites.
Accelerated stochastic approximation
B. Delyon and A. Juditsky · 1993
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Nicol N. Schraudolph · 1999
Earlier work this paper cites.
Probability and Random Processes
Geoffrey Grimmett and David Stirzaker · 2001
Earlier work this paper cites.
Unicorns do exist: A tutorial on “proving” the null hypothesis
David L Streiner · 2003
Cited alongside, same era.
Introductory Lecture on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Cited alongside, same era.
Testing Statistical Hypotheses
Erich L. Lehmann and Joseph P. Romano · 2005
Cited alongside, same era.
Fixed-width output analysis for markov chain monte carlo
Galin L Jones, Murali Haran, Brian S Caffo, and Ronald Neath · 2006
Cited alongside, same era.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Cited alongside, same era.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Deep residual networks for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
Atilim Günes Baydin, Robert Cornish, David Martínez Rubio, Mark Schmidt, and Frank Wood · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Batch means and spectral variance estimators in markov chain monte carlo
James M Flegal and Galin L Jones · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Online algorithms and stochastic approximations
Léon Bottou · 2012
Cited alongside, same era.
Tuning-free step-size adaption
Ashique Rupam Mahmood, Richard S. Sutton, Thomas Degris, and Patrick M. Pilarski · 2012
Cited alongside, same era.
Introduction to Linear Regression Analysis
D.C. Montgomery, E.A. Peck, and G.G. Vining · 2012
Cited alongside, same era.
Later among the works it cites.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Later among the works it cites.
Understanding the role momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Later among the works it cites.
Using statistics to automate stochastic optimization
Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Later among the works it cites.
Pytorch cifar
Kuang Liu · 2019
Later among the works it cites.
Quasi-hyperbolic momentum and adam for deep learning
Jerry Ma and Denis Yarats · 2019
Later among the works it cites.
Pytorch word language model
PyTorch · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Sharan Vaswani, Aaron Mishkin, Issam Laradji, Mark Schmidt, Gauthier Gidel, and Simon Lacoste-Julian · 2019
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2019
Later among the works it cites.