Fetching the paper…
Reading the bibliography…
Motivated by the recent successes of neural networks that have the ability to fit the data perfectly and generalize well, we study the noiseless model in the fundamental least-squares setup.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Incremental gradient algorithms with stepsizes bounded away from zero
M. V. Solodov · 1998
Earlier work this paper cites.
An incremental gradient(-projection) method with momentum term and adaptive stepsize rule
P. Tseng · 1998
Earlier work this paper cites.
Learning with Kernels
B. Schölkopf and A. J. Smola · 2002
Earlier work this paper cites.
Kernel Methods for Pattern Analysis
J. Shawe-Taylor and N. Cristianini · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Earlier work this paper cites.
Optimal rates for regularized least squares regression
I. Steinwart, D. R. Hush, and C. Scovel · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Earlier work this paper cites.
A lecture on the averaging process
D. Aldous and D. Lanoue · 2012
Earlier work this paper cites.
Open problem: is averaging needed for strongly convex stochastic gradient descent?
O. Shamir · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate o(1/n)
F. Bach and E. Moulines · 2013
Cited alongside, same era.
Fast convergence of stochastic gradient descent under a strong growth condition
M. Schmidt and N. L. Roux · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
Online learning as stochastic approximation of regularization paths: optimality and almost-sure convergence
P. Tarres and Y. Yao · 2014
Cited alongside, same era.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
L. Pillaud-Vivien, A. Rudi, and F. Bach · 2018
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Later among the works it cites.
The step decay schedule: a near optimal, geometrically decaying learning rate procedure for least squares
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli · 2019
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
N. J. A. Harvey, C. Liaw, Y. Plan, and S. Randhawa · 2019
Later among the works it cites.
Making the last iterate of sgd information theoretically optimal
P. Jain, D. Nagaraj, and P. Netrapalli · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Dieuleveut and F. Bach · 2016
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression
A. Dieuleveut, N. Flammarion, and F. Bach · 2017
Cited alongside, same era.
Stochastic composite least-squares regression with convergence rate o ( 1 / n ) o(1/n)
N. Flammarion and F. Bach · 2017
Cited alongside, same era.
On exponential convergence of sgd in non-convex over-parametrized learning
R. Bassily, M. Belkin, and S. Ma · 2018
Cited alongside, same era.
Accelerating stochastic gradient descent for least squares regression
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
S. Ma, R. Bassily, and M. Belkin · 2018
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
S. Vaswani, F. Bach, and M. Schmidt
Cited in the paper.
K. Jun, A. Cutkosky, and F. Orabona · 2019
Later among the works it cites.
Tight nonparametric convergence rates for stochastic gradient descent under the noiseless linear model
R. Berthier, F. Bach, and P. Gaillard · 2020
Later among the works it cites.
Accelerating sgd with momentum for over-parameterized learning
C. Liu and M. Belkin · 2020
Later among the works it cites.
Fast and furious convergence: stochastic second order methods under interpolation
S. Y. Meng, S. Vaswani, I. H. Laradji, M. Schmidt, and S. Lacoste-Julien · 2020
Later among the works it cites.
Stochastic gradient descent in Hilbert scales: smoothness, preconditioning and earlier stopping
N. Mücke and E. Reiss · 2020
Later among the works it cites.
Benign overfitting of constant-stepsize sgd for linear regression
D. Zou, J. Wu, V. Braverman, Q. Gu, and S. M. Kakade · 2021
Closest in time.