Fetching the paper…
Reading the bibliography…
Tuning the step size of stochastic gradient descent is tedious and error prone.
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays
U. Alon, N. Barkai, D. A. Notterman, K. Gish, S. Ybarra, D. Mack, and A. J. Levine · 1999
Earlier work this paper cites.
Learning additive models online with fast evaluating kernels
M. Herbster · 2001
Earlier work this paper cites.
Non-projective dependency parsing using spanning tree algorithms
R. McDonald, F. Pereira, K. Ribarov, and J. Hajic · 2005
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
J. Blitzer, R. McDonald, and F. Pereira · 2006
Earlier work this paper cites.
Online passive-aggressive algorithms
K. Crammer, O. Dekel, J. Keshet, S. Shalev-Shwartz, and Y. Singer · 2006
Earlier work this paper cites.
Detection of non-coding rnas on the basis of predicted secondary structure formation free energy change
A. V. Uzilov, J. M. Keegan, and D. H. Mathews · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Large scale online learning of image similarity through ranking
G. Chechik, V. Sharma, U. Shalit, and S. Bengio · 2010
Earlier work this paper cites.
Mnist handwritten digit database. at&t labs, 2010
Y. LeCun, C. Cortes, and C. Burges · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2013
Earlier work this paper cites.
Online passive-aggressive algorithms for non-negative matrix factorization and completion
M. Blondel, Y. Kubo, and N. Ueda · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
UCI machine learning repository, 2017
D. Dua and C. Graff · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
L. Balles and P. Hennig · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
S. Ma, R. Bassily, and M. Belkin · 2018
Cited alongside, same era.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien · 2019
Later among the works it cites.
Training neural networks for and by interpolation
L. Berrada, A. Zisserman, and M. P. Kumar · 2020
Later among the works it cites.
Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation
R. M. Gower, O. Sebbouh, and N. Loizou · 2020
Later among the works it cites.
Unified Analysis of Stochastic Gradient Methods for Composite Convex and Smooth Optimization
A. Khaled, O. Sebbouh, N. Loizou, R. M. Gower, and P. Richtárik · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
T. Liang and A. Rakhlin · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
The importance of better models in stochastic optimization
H. Asi and J. C. Duchi · 2019
Cited alongside, same era.
Does data interpolation contradict statistical optimality?
M. Belkin, A. Rakhlin, and A. B. Tsybakov · 2019
Cited alongside, same era.
Sgd: General analysis and improved rates
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Cited alongside, same era.
A modern introduction to online learning
F. Orabona · 2019
Cited alongside, same era.
Later among the works it cites.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
N. Loizou, S. Vaswani, I. Laradji, and S. Lacoste-Julien · 2020
Later among the works it cites.
R. Yuan, A. Lazaric, and R. M. Gower · 2020
Later among the works it cites.
Comment on stochastic polyak step-size: Performance of ALI-G
L. Berrada, A. Zisserman, and M. P. Kumar · 2021
Later among the works it cites.
Accelerated, optimal, and parallel: Some results on model-based stochastic optimization
K. N. Chadha, G. Cheng, and J. C. Duchi · 2021
Later among the works it cites.
Fragments d’optimisation différentiable-théories et algorithmes
J. C. Gilbert · 2021
Later among the works it cites.
Stochastic polyak stepsize with a moving target
R. M. Gower, A. Defazio, and M. Rabbat · 2021
Later among the works it cites.
An online passive-aggressive algorithm for difference-of-squares classification
L. K. Saul · 2021
Later among the works it cites.