Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) has been shown to generalize well in many deep learning applications.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B., Smola, A. J., Bach, F., et al · 2002
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Ghadimi, S. and Lan, G · 2012
Earlier work this paper cites.
Lacoste-Julien, S., Schmidt, M., and Bach, F · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o ( 1 / n ) o(1/n)
Bach, F. and Moulines, E · 2013
Earlier work this paper cites.
A risk comparison of ordinary least squares vs ridge regression
Dhillon, P. S., Foster, D. P., Kakade, S. M., and Ungar, L. H · 2013
Earlier work this paper cites.
Theory of convex optimization for machine learning
Bubeck, S · 2014
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Hazan, E. and Kale, S · 2014
Earlier work this paper cites.
Random design analysis of ridge regression
Hsu, D. J., Kakade, S. M., and Zhang, T · 2014
Cited alongside, same era.
Non-parametric stochastic approximation with large step sizes
Dieuleveut, A. and Bach, F. R · 2015
Cited alongside, same era.
Deep residual learning for image recognition. corr abs/1512.03385 (2015), 2015
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression
Dieuleveut, A., Flammarion, N., and Bach, F · 2017
Cited alongside, same era.
Optimal rates for multi-pass stochastic gradient methods
Lin, J. and Rosasco, L · 2017
Cited alongside, same era.
Symmetric multivariate and related distributions
Fang, K.-T., Kotz, S., and Ng, K. W · 2018
A generic acceleration framework for stochastic composite optimization
Kulunchakov, A. and Mairal, J · 2019
Later among the works it cites.
Beating sgd saturation with tail-averaging and minibatching
Mücke, N., Neu, G., and Rosasco, L · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2020
Later among the works it cites.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 2020
Later among the works it cites.
Berthier, R., Bach, F., and Gaillard, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Iterate averaging as regularization for stochastic gradient descent
Neu, G. and Rosasco, L · 2018
Cited alongside, same era.
A universally optimal multistage accelerated stochastic gradient method
Aybat, N. S., Fallah, A., Gurbuzbalaban, M., and Ozdaglar, A · 2019
Cited alongside, same era.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
Davis, D., Drusvyatskiy, D., and Charisopoulos, V · 2019
Cited alongside, same era.
Ge, R., Kakade, S. M., Kidambi, R., and Netrapalli, P · 2019
Cited alongside, same era.
Jain, P., Kakade, S. M., Kidambi, R., Netrapalli, P., Pillutla, V. K., and Sidford, A
Cited in the paper.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Jain, P., Netrapalli, P., Kakade, S. M., Kidambi, R., and Sidford, A
Cited in the paper.
Tsigler, A. and Bartlett, P. L · 2020
Later among the works it cites.
Pan, R., Ye, H., and Zhang, T · 2021
Closest in time.
Last iterate convergence of sgd for least-squares in the interpolation regime
Varre, A., Pillaud-Vivien, L., and Flammarion, N · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Closest in time.