Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) has emerged as the quintessential method in a data scientist's toolbox.
Herbert Robbins and Sutton Monro, A stochastic approximation method , The Annals of Mathematical Statistics 22
1951
Earlier work this paper cites.
Kai Lai Chung, On a stochastic approximation method , The Annals of Mathematical Statistics (1954), 463–483
1954
Earlier work this paper cites.
Jerome Sacks, Asymptotic distribution of stochastic approximation procedures , The Annals of Mathematical Statistics 29
1958
Earlier work this paper cites.
Vaclav Fabian, On asymptotic normality in stochastic approximation , The Annals of Mathematical Statistics 39
1968
Earlier work this paper cites.
Erwin Bolthausen, Exact convergence rates in some martingale central limit theorems , The Annals of Probability (1982), 672–688
1982
Earlier work this paper cites.
Peter J Bickel and David A Freedman, Bootstrapping regression models with many parameters , Festschrift for Erich L. Lehmann (1983), 28–48
1983
Earlier work this paper cites.
Stephen Portnoy, Asymptotic behavior of m m estimators of p p regression parameters when p 2 / n p^{2}/n is large; ii. normal approximation , The Annals of Statistics 13
1985
Earlier work this paper cites.
by same author, Asymptotic behavior of the empirical distribution of M-estimated residuals from a regression model with many parameters , The Annals of Statistics 14
1986
Earlier work this paper cites.
Erich Haeusler, On the rate of convergence in the central limit theorem for martingales with discrete and continuous time , The Annals of Probability (1988), 275–299
1988
Earlier work this paper cites.
David Ruppert, Efficient estimations from a slowly convergent Robbins-Monro process , Tech. report, Cornell University Operations Research and Industrial Engineering, 1988
1988
Earlier work this paper cites.
Enno Mammen, Asymptotics with increasing dimension for robust regression with applications to the bootstrap , The Annals of Statistics (1989), 382–400
1989
Earlier work this paper cites.
Alexander Shapiro, Asymptotic properties of statistical estimators in stochastic programming , The Annals of Statistics 17
1989
Earlier work this paper cites.
Peter W Glynn and Ward Whitt, Estimating the asymptotic variance with batch means , Operations Research Letters 10
1991
Earlier work this paper cites.
Boris T Polyak and Anatoli B Juditsky, Acceleration of stochastic approximation by averaging , SIAM journal on control and optimization 30
1992
Earlier work this paper cites.
by same author, Bootstrap and wild bootstrap for high dimensional linear models , The annals of statistics 21
1993
Earlier work this paper cites.
J Harold, G Kushner, and George Yin, Stochastic approximation and recursive algorithm and applications , Application of Mathematics 35
1997
Earlier work this paper cites.
Dimitris N Politis, Joseph P Romano, and Michael Wolf, Subsampling , Springer Science & Business Media, 1999
1999
Earlier work this paper cites.
Vivek S Borkar, Stochastic approximation: a dynamical systems viewpoint , vol. 48, Springer, 2009
2009
Earlier work this paper cites.
James M Flegal and Galin L Jones, Batch means and spectral variance estimators in Markov chain Monte Carlo , The Annals of Statistics 38
2010
Earlier work this paper cites.
Albert Benveniste, Michel Métivier, and Pierre Priouret, Adaptive algorithms and stochastic approximations , vol. 22, Springer Science & Business Media, 2012
2012
Earlier work this paper cites.
Young Min Kim, Soumendra N Lahiri, and Daniel J Nordman, A progressive block empirical likelihood method for time series , Journal of the American Statistical Association 108
2013
Earlier work this paper cites.
SN Lahiri, Resampling methods for dependent data , Springer Science & Business Media, 2013
2013
Earlier work this paper cites.
Jean-Christophe Mourrat, On the rate of convergence in the martingale central limit theorem , Bernoulli (2013), 633–645
2013
Earlier work this paper cites.
Sébastien Bubeck, Convex optimization: Algorithms and complexity , Foundations and Trends® in Machine Learning 8
2015
Earlier work this paper cites.
Ethan X Fang, Yang Ning, and Han Liu, Testing and confidence intervals for high dimensional proportional hazards models , Journal of the Royal Statistical Society. Series B (Statistical Methodology) (2017), 1415–1437
2017
Earlier work this paper cites.
RE Gaunt, A Pickett, and G Reinert, Chi-square approximation by Stein’s method with application to Pearson’s statistic , Annals of Applied Probability 27
2017
Earlier work this paper cites.
Panos Toulis and Edoardo M Airoldi, Asymptotic and finite-sample properties of estimators based on stochastic gradients , The Annals of Statistics 45
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
Léon Bottou, Frank E Curtis, and Jorge Nocedal, Optimization methods for large-scale machine learning , SIAM review 60
2018
Cited alongside, same era.
Noureddine El Karoui and Elizabeth Purdom, Can we trust the bootstrap in high-dimensions? the case of linear models , The Journal of Machine Learning Research 19
2018
Cited alongside, same era.
Yixin Fang, Jinfeng Xu, and Lei Yang, Online bootstrap confidence intervals for the stochastic gradient descent estimator , The Journal of Machine Learning Research 19
2018
Cited alongside, same era.
2021
Later among the works it cites.
Robert Lunde, Purnamrita Sarkar, and Rachel Ward, Bootstrapping the error of Oja’s algorithm , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Courtney Paquette, Kiwon Lee, Fabian Pedregosa, and Elliot Paquette, SGD in the large: Average-case analysis, asymptotics, and stepsize criticality , Conference on Learning Theory, PMLR, 2021, pp. 3548–3626
2021
Later among the works it cites.
Courtney Paquette and Elliot Paquette, Dynamics of stochastic momentum methods on large-scale, quadratic models , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Tianyang Li, Liu Liu, Anastasios Kyrillidis, and Constantine Caramanis, Statistical inference using SGD , Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, 2018
2018
Cited alongside, same era.
Adrian Röllin, On quantitative bounds in the mean martingale central limit theorem , Statistics & Probability Letters 138
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , vol. 47, Cambridge university press, 2018
2018
Cited alongside, same era.
Andreas Anastasiou, Krishnakumar Balasubramanian, and Murat A Erdogdu, Normal approximation for stochastic gradient descent via non-asymptotic rates of martingale CLT , Conference on Learning Theory, PMLR, 2019
2019
Cited alongside, same era.
Hilal Asi and John C Duchi, Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity , SIAM Journal on Optimization (2019)
2019
Cited alongside, same era.
Othmane Sebbouh, Robert M Gower, and Aaron Defazio, Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball , Conference on Learning Theory, PMLR, 2021, pp. 3935–3971
2021
Later among the works it cites.
Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion, Last iterate convergence of SGD for Least-Squares in the Interpolation regime , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Lu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, and Murat A Erdogdu, An analysis of constant step size SGD in the non-convex regime: Asymptotic normality and bias , Advances in Neural Information Processing (2021)
2021
Later among the works it cites.
Wanrong Zhu, Xi Chen, and Wei Biao Wu, Online covariance matrix estimation in stochastic gradient descent , Journal of the American Statistical Association (2021), 1–12
2021
Later among the works it cites.
2021
Later among the works it cites.
by same author, High-dimensional limit theorems for SGD: Effective dynamics and critical scaling , Advances in Neural Information Processing Systems, 2022
2022
Later among the works it cites.
Kabir Aladin Chandrasekher, Ashwin Pananjady, and Christos Thrampoulidis, Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective , The Annals of Statistics (to appear) (2022)
2022
Later among the works it cites.
Xi Chen, Qiang Liu, and Xin T Tong, Dimension independent excess risk by stochastic gradient descent , Electronic Journal of Statistics 16
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
De Huang, Jonathan Niles-Weed, Joel A Tropp, and Rachel Ward, Matrix concentration for products , Foundations of Computational Mathematics 22
2022
Later among the works it cites.
Sokbae Lee, Yuan Liao, Myung Hwan Seo, and Youngki Shin, Fast and robust online inference with stochastic gradient descent via random scaling , Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, 2022, pp. 7381–7389
2022
Later among the works it cites.
Jun Liu and Ye Yuan, On almost sure convergence rates of stochastic gradient methods , Conference on Learning Theory, PMLR, 2022, pp. 2963–2983
2022
Later among the works it cites.
2022
Later among the works it cites.
by same author, Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions , Advances in Neural Information Processing Systems, 2022
2022
Later among the works it cites.
Qi-Man Shao and Zhuo-Song Zhang, Berry–Esseen bounds for multivariate nonlinear statistics with applications to M-estimators and stochastic gradient descent algorithms , Bernoulli 28
2022
Later among the works it cites.
Jingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu, and Sham Kakade, Last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression , International Conference on Machine Learning, PMLR, 2022, pp. 24280–24314
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Jerry Chee, Hwanwoo Kim, and Panos Toulis, “Plus/minus the learning rate”: Easy and Scalable Statistical Inference with SGD , 26th International Conference on Artificial Intelligence and Statistics (AISTATS), 2023
2023
Closest in time.
2023
Closest in time.