Fetching the paper…
Reading the bibliography…
We investigate the stochastic gradient descent (SGD) method where the step size lies within a banded region instead of being given by a fixed formula.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
Kai Lai Chung · 1954
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Lennart Ljung · 1977
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Optimal stochastic search and adaptive momentum
Todd K Leen and Genevieve B Orr · 1994
Earlier work this paper cites.
Two approaches to optimal annealing
Todd K Leen, Bernhard Schottky, and David Saad · 1998
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter L Bartlett, and Pradeep K Ravikumar · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp, COURSERA: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Later among the works it cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Later among the works it cites.
How to make the gradients small stochastically: Even faster convex and nonconvex SGD
Zeyuan Allen-Zhu · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Random Barzilai-Borwein step size for mini-batch algorithms
Zhuang Yang, Cheng Wang, Zhemin Zhang, and Jonathan Li · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Escaping from saddle points-online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
A nonmonotone learning rate strategy for SGD training of deep neural networks
Nitish Shirish Keskar and George Saon · 2015
Cited alongside, same era.
Adam: A method for stochastic gradient descent
Diederik P Kingma and Jimmy Lei Ba · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Barzilai-Borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Cited alongside, same era.
Exponential decay sine wave learning rate for fast deep neural network training
Wangpeng An, Haoqian Wang, Yulun Zhang, and Qionghai Dai · 2017
Cited alongside, same era.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Later among the works it cites.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Later among the works it cites.
Making the last iterate of SGD information theoretically optimal
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Sharan Vaswani, Aaron Mishkin, Issam Laradji, Mark Schmidt, Gauthier Gidel, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Stochastic quasi-gradient methods: Variance reduction via Jacobian sketching
Robert M Gower, Peter Richtárik, and Francis Bach · 2020
Later among the works it cites.
Exponential step sizes for non-convex optimization
Xiaoyu Li, Zhenxun Zhuang, and Francesco Orabona · 2020
Later among the works it cites.
Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Laradji, and Simon Lacoste-Julien · 2020
Later among the works it cites.