Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) has achieved great success due to its superior performance in both optimization and generalization.
Stability of randomized learning algorithms
Elisseeff, A · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L · 2007
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y · 2007
Earlier work this paper cites.
Benign overfitting in ridge regression
Tsigler, A · 2009
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o ( 1 / n ) o(1/n)
Bach, F · 2013
Earlier work this paper cites.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
Raskutti, G · 2014
Earlier work this paper cites.
Learning with incremental iterative regularization
Rosasco, L · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M · 2016
Earlier work this paper cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Dieuleveut, A · 2017
Earlier work this paper cites.
Optimal rates for multi-pass stochastic gradient methods
Lin, J · 2017
Earlier work this paper cites.
Early stopping for kernel boosting algorithms: A general analysis with localized complexities
Wei, Y · 2017
Earlier work this paper cites.
On exponential convergence of sgd in non-convex over-parametrized learning
Bassily, R · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L · 2018
Cited alongside, same era.
Stability and convergence trade-off of iterative optimization algorithms
Chen, Y · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Dobriban, E · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S · 2018
Cited alongside, same era.
Data-dependent stability of stochastic gradient descent
Random shuffling beats sgd after finite epochs
Haochen, J · 2019
Later among the works it cites.
Beating sgd saturation with tail-averaging and minibatching
Mücke, N · 2019
Later among the works it cites.
On the number of variables to use in principal component regression
Xu, J · 2019
Later among the works it cites.
Sgd with shuffling: optimal rates without component convexity and large epoch requirements
Ahn, K · 2020
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L · 2020
Later among the works it cites.
Stability of stochastic gradient descent on nonsmooth convex losses
Bassily, R · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuzborskij, I · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Ma, S · 2018
Cited alongside, same era.
Iterate averaging as regularization for stochastic gradient descent
Neu, G · 2018
Cited alongside, same era.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Pillaud-Vivien, L · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Ali, A · 2019
Cited alongside, same era.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Ge, R · 2019
Cited alongside, same era.
Jain, P
Cited in the paper.
How good is sgd with random shuffling?
Safran, I · 2020
Later among the works it cites.
Generalization performance of multi-pass stochastic gradient descent with convex loss functions
Lei, Y · 2021
Later among the works it cites.
Last iterate risk bounds of sgd with decaying stepsize for overparameterized linear regression
Wu, J · 2021
Later among the works it cites.
Stability of sgd: Tightness analysis and improved bounds
Zhang, Y · 2021
Later among the works it cites.