Fetching the paper…
Reading the bibliography…
In this work, we propose new adaptive step size strategies that improve several stochastic gradient methods.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Introduction to optimization. optimization software
Boris T. Polyak · 1987
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Subgradient methods
Stephen Boyd, Lin Xiao, and Almir Mutapcic · 2003
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis Bach · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Zheng Qu, Peter Richtárik, and Tong Zhang · 2015
Earlier work this paper cites.
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2016
Cited alongside, same era.
Barzilai-Borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Stochastic gradient descent with Polyak’s learning rate
Adam M. Oberman and Mariana Prazeres · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konečný, Martin Jaggi, Virginia Smith, Michael I. Jordan, Peter Richtárik, and Martin Takáč · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
SGD and Hogwild! convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takác · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Later among the works it cites.
Stochastic first-order methods: non-asymptotic and computer-aided analyses via potential functions
Adrien Taylor and Francis Bach · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
Training neural networks for and by interpolation
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2020
Later among the works it cites.
BackPACK: Packing more into backprop
Felix Dangel, Frederik Kunstner, and Philipp Hennig · 2020
Later among the works it cites.
AdaScale SGD: A user-friendly algorithm for distributed training
Tyler Johnson, Pulkit Agrawal, Haijie Gu, and Carlos Guestrin · 2020
Later among the works it cites.
Don’t jump through hoops and remove those loops: SVRG and katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horváth, and Peter Richtárik · 2020
Later among the works it cites.
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2020
Later among the works it cites.
Stochastic reformulations of linear systems: algorithms and convergence theory
Peter Richtárik and Martin Takáč · 2020
Later among the works it cites.
Stochastic Polyak stepsize with a moving target
Robert M. Gower, Aaron Defazio, and Michael Rabbat · 2021
Later among the works it cites.
Stochastic Polyak step-size for SGD: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Hadj Laradji, and Simon Lacoste-Julien · 2021
Later among the works it cites.
Cutting some slack for SGD with adaptive Polyak stepsizes
Robert M. Gower, Mathieu Blondel, Nidham Gazagnadou, and Fabian Pedregosa · 2022
Closest in time.
SP2: A second order stochastic Polyak method
Shuang Li, William J. Swartworth, Martin Takáč, Deanna Needell, and Robert M. Gower · 2022
Closest in time.