Fetching the paper…
Reading the bibliography…
Here we develop variants of SGD (stochastic gradient descent) with an adaptive step size that make use of the sampled loss values.
“A Modern Introduction to Online Learning”, 2019
Francesco Orabona · 1912
Earlier work this paper cites.
“Introduction to Optimization. Translations series in mathematics and engineering”
B.T. Polyak · 1987
Earlier work this paper cites.
“Fundamentals of convex analysis” Abridged version of ıt Convex analysis and minimization algorithms. I [Springer, Berlin, 1993; MR1261420 (95m:90001)] and ıt II [ibid.; MR1295240 (95m:90002)], Grundlehren Text Editions
Jean-Baptiste Hiriart-Urruty and Claude Lemar“’echal · 2001
Earlier work this paper cites.
“Numerical optimization”, Springer Series in Operations Research and Financial Engineering
Jorge Nocedal and Stephen. Wright · 2006
Earlier work this paper cites.
“Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers.”
Stephen. Boyd et al · 2011
Earlier work this paper cites.
“LIBSVM: a library for support vector machines”
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
“Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”
John Duchi, Elad Hazan and Yoram Singer · 2011
Earlier work this paper cites.
“Introductory Lectures on Convex Optimization: A Basic Course”
Y. Nesterov · 2013
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“First-order methods in optimization” 25
Amir Beck · 2017
Earlier work this paper cites.
“A Sampling Kaczmarz–Motzkin Algorithm for Linear Feasibility”
Jes“’us. De, Jamie Haddock and Deanna Needell · 2017
Cited alongside, same era.
“Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron”
Sharan Vaswani, Francis Bach and Mark Schmidt · 2018
Cited alongside, same era.
“The importance of better models in stochastic optimization”
Hilal Asi and John. Duchi · 2019
Cited alongside, same era.
“Stochastic model-based minimization of weakly convex functions”
Damek Davis and Dmitriy Drusvyatskiy · 2019
Cited alongside, same era.
“Efficiency of minimizing compositions of convex functions and smooth maps”
D. Drusvyatskiy and C. Paquette · 2019
Cited alongside, same era.
“Revisiting the Polyak step size”
Elad Hazan and Sham Kakade · 2019
“Better Theory for SGD in the Nonconvex World”
Ahmed Khaled and Peter Richtarik · 2020
Later among the works it cites.
“Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence”
Nicolas Loizou, Sharan Vaswani, Issam Laradji and Simon Lacoste-Julien · 2020
Later among the works it cites.
“Adaptive Gradient Descent without Descent”
Yura Malitsky and Konstantin Mishchenko · 2020
Later among the works it cites.
“Accelerated, Optimal, and Parallel: Some Results on Model-Based Stochastic Optimization”
Karan. Chadha, Gary Cheng and John. Duchi · 2021
Later among the works it cites.
“Stochastic Polyak Stepsize with a Moving Target”, 2021
Robert. Gower, Aaron Defazio and Mike Rabbat · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates”
Sharan Vaswani et al · 2019
Cited alongside, same era.
“Training Neural Networks for and by Interpolation”
Leonard Berrada, Andrew Zisserman and M. Kumar · 2020
Cited alongside, same era.
“SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation”
Robert. Gower, Othmane Sebbouh and Nicolas Loizou · 2020
Cited alongside, same era.
Later among the works it cites.
“Cutting Some Slack for SGD with Adaptive Polyak Stepsizes”
Robert. Gower, Mathieu Blondel, Nidham Gazagnadou and Fabian Pedregosa · 2022
Later among the works it cites.
“A Stochastic Bundle Method for Interpolating Networks”
Alasdair Paren, Leonard Berrada, Rudra P.. Poudel and M. Kumar · 2022
Later among the works it cites.
“Handbook of Convergence Theorems for (Stochastic) Gradient Methods”
Guillaume Garrigos and Robert. Gower · 2023
Closest in time.
“A Model-Based Method for Minimizing CVaR and Beyond”
Si Meng and Robert. Gower · 2023
Closest in time.