Fetching the paper…
Reading the bibliography…
Recently the "SP" (Stochastic Polyak step size) method has emerged as a competitive adaptive method for setting the step sizes of SGD.
Introduction to optimization
B. Polyak · 1987
Earlier work this paper cites.
Automatic Hessians by reverse accumulation
B. Christianson · 1992
Earlier work this paper cites.
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays
U. Alon, N. Barkai, D. A. Notterman, K. Gish, S. Ybarra, D. Mack, and A. J. Levine · 1999
Earlier work this paper cites.
Numerical optimization
S. Wright and J. Nocedal · 1999
Earlier work this paper cites.
Predicting the clinical status of human breast cancer by using gene expression profiles
M. West, C. Blanchette, H. Dressman, E. Huang, S. Ishida, R. Spang, H. Zuzan, J. A. Olson, J. R. Marks, and J. R. Nevins · 2001
Earlier work this paper cites.
Online passive-aggressive algorithms
K. Crammer, O. Dekel, J. Keshet, S. Shalev-Shwartz, and Y. Singer · 2006
Earlier work this paper cites.
Fast algorithms for projection on an ellipsoid
Y.-H. Dai · 2006
Earlier work this paper cites.
On the use of stochastic Hessian information in optimization methods for machine learning
R. H. Byrd, G. M. Chin, W. Neveitt, and J. Nocedal · 2011
Earlier work this paper cites.
A literature survey of benchmark functions for global optimization problems
M. Jamil and X.-S. Yang · 2013
Earlier work this paper cites.
Convergence rates of sub-sampled Newton methods
M. A. Erdogdu and A. Montanari · 2015
Earlier work this paper cites.
Global convergence of online limited memory BFGS
A. Mokhtari and A. Ribeiro · 2015
Earlier work this paper cites.
A multi-batch l-bfgs method for machine learning
A. S. Berahas, J. Nocedal, and M. Takáč · 2016
Earlier work this paper cites.
Stochastic block BFGS: Squeezing more curvature out of data
R. M. Gower, D. Goldfarb, and P. Richtárik · 2016
Earlier work this paper cites.
Provable efficient online matrix completion via non-convex stochastic gradient descent
C. Jin, S. M. Kakade, and P. Netrapalli · 2016
Cited alongside, same era.
A linearly-convergent stochastic L-BFGS algorithm
P. Moritz, R. Nishihara, and M. I. Jordan · 2016
Cited alongside, same era.
SDNA: Stochastic dual Newton ascent for empirical risk minimization
Z. Qu, P. Richtárik, M. Takáč, and O. Fercoq · 2016
Cited alongside, same era.
A superlinearly-convergent proximal newton-type method for the optimization of finite sums
A. Rodomanov and D. Kropotov · 2016
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
N. Agarwal, B. Bullins, and E. Hazan · 2017
Cited alongside, same era.
Distributed restarting newtoncg method for large-scale empirical risk minimization
M. Jahani, X. He, C. Ma, D. Mudigere, A. Mokhtari, A. Ribeiro, and M. Takac · 2017
Stochastic cubic regularization for fast nonconvex optimization
N. Tripuraneni, M. Stern, C. Jin, J. Regier, and M. I. Jordan · 2018
Later among the works it cites.
Near-optimal methods for minimizing star-convex functions and beyond
O. Hinder, A. Sidford, and N. S. Sohoni · 2019
Later among the works it cites.
Stochastic Newton and cubic Newton methods with simple local linear-quadratic rates
D. Kovalev, K. Mishchenko, and P. Richtarik · 2019
Later among the works it cites.
Sub-sampled newton methods
F. Roosta-Khorasani and M. W. Mahoney · 2019
Later among the works it cites.
Training neural networks for and by interpolation
L. Berrada, A. Zisserman, and M. P. Kumar · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sub-sampled cubic regularization for non-convex optimization
J. M. Kohler and A. Lucchi · 2017
Cited alongside, same era.
General heuristics for nonconvex quadratically constrained quadratic programming
J. Park and S. Boyd · 2017
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
M. Pilanci and M. J. Wainwright · 2017
Cited alongside, same era.
Stochastic quasi-newton methods for nonconvex stochastic optimization, 2017
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Cited alongside, same era.
Exact and inexact subsampled Newton methods for optimization
R. Bollapragada, R. H. Byrd, and J. Nocedal · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
S. Ma, R. Bassily, and M. Belkin · 2018
Cited alongside, same era.
R. M. Gower, O. Sebbouh, and N. Loizou · 2020
Later among the works it cites.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
N. Loizou, S. Vaswani, I. Laradji, and S. Lacoste-Julien · 2020
Later among the works it cites.
An algorithm for projecting a point onto a level set of a quadratic function
W. Sosa and F. MP Raupp · 2020
Later among the works it cites.
San: Stochastic average newton algorithm for minimizing finite sums, 2021
J. Chen, R. Yuan, G. Garrigos, and R. M. Gower · 2021
Later among the works it cites.
Stochastic polyak stepsize with a moving target
R. M. Gower, A. Defazio, and M. Rabbat · 2021
Later among the works it cites.
Convergence of newton-mr under inexact hessian information
Y. Liu and F. Roosta · 2021
Later among the works it cites.
Cutting some slack for sgd with adaptive polyak stepsizes
R. M. Gower, M. Blondel, N. Gazagnadou, and F. Pedregosa · 2022
Closest in time.