Fetching the paper…
Reading the bibliography…
Structured non-convex learning problems, for which critical points have favorable statistical properties, arise frequently in statistical machine learning.
Herbert Robbins and Sutton Monro, A stochastic approximation method , The annals of mathematical statistics (1951), 400–407
1951
Earlier work this paper cites.
Kai Lai Chung, On a stochastic approximation method , The Annals of Mathematical Statistics (1954), 463–483
1954
Earlier work this paper cites.
Jerome Sacks, Asymptotic distribution of stochastic approximation procedures , The Annals of Mathematical Statistics 29
1958
Earlier work this paper cites.
Vaclav Fabian, On asymptotic normality in stochastic approximation , The Annals of Mathematical Statistics 39
1968
Earlier work this paper cites.
DA Freedman and P Diaconis, On inconsistent m m -estimators , The Annals of Statistics 10
1982
Earlier work this paper cites.
Andrew Blake and Andrew Zisserman, Visual reconstruction , 1987
1987
Earlier work this paper cites.
Yuri Kifer, Random perturbations of dynamical systems , Nonlinear Problems in Future Particle Accelerators. World Scientific (1988), 189
1988
Earlier work this paper cites.
David Ruppert, Efficient estimations from a slowly convergent robbins-monro process , Tech. report, Cornell University Operations Research and Industrial Engineering, 1988
1988
Earlier work this paper cites.
Alexander Shapiro, Asymptotic properties of statistical estimators in stochastic programming , The Annals of Statistics 17
1989
Earlier work this paper cites.
Boris T Polyak and Anatoli B Juditsky, Acceleration of stochastic approximation by averaging , SIAM Journal on Control and Optimization 30
1992
Earlier work this paper cites.
Michel Benaim, A dynamical system approach to stochastic approximations , SIAM Journal on Control and Optimization 34
1996
Earlier work this paper cites.
P Priouret and A Yu Veretenikov, A remark on the stability of the lms tracking algorithm , Stochastic analysis and applications 16
1998
Earlier work this paper cites.
Jean-Claude Fort and Gilles Pages, Asymptotic behavior of a markovian stochastic algorithm with constant step , SIAM journal on control and optimization 37
1999
Earlier work this paper cites.
Rafik Aguech, Eric Moulines, and Pierre Priouret, On a perturbation approach for the analysis of stochastic tracking algorithms , SIAM Journal on Control and Optimization 39
2000
Earlier work this paper cites.
Jonathan C Mattingly, Andrew M Stuart, and Desmond J Higham, Ergodicity for sdes and approximations: locally lipschitz vector fields and degenerate noise , Stochastic processes and their applications 101
2002
Earlier work this paper cites.
Peter J Huber, Robust statistics , vol. 523, John Wiley & Sons, 2004
2004
Earlier work this paper cites.
Yurii Nesterov and Boris T Polyak, Cubic regularization of newton method and its global performance , Mathematical Programming 108
2006
Earlier work this paper cites.
Charles J Geyer, On the asymptotics of constrained m m -estimation , The Annals of Statistics 22
2010
Earlier work this paper cites.
Eric Moulines and Francis R Bach, Non-asymptotic analysis of stochastic approximation algorithms for machine learning , Advances in Neural Information Processing Systems, 2011, pp. 451–459
2011
Earlier work this paper cites.
Sean P Meyn and Richard L Tweedie, Markov chains and stochastic stability , Springer Science & Business Media, 2012
2012
Earlier work this paper cites.
Saeed Ghadimi and Guanghui Lan, Stochastic first-and zeroth-order methods for nonconvex stochastic programming , SIAM Journal on Optimization 23
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Deanna Needell, Rachel Ward, and Nati Srebro, Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm , Advances in neural information processing systems, 2014, pp. 1017–1025
2014
Earlier work this paper cites.
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan, Escaping from saddle points - online stochastic gradient for tensor decomposition , Conference on Learning Theory, 2015, pp. 797–842
2015
Earlier work this paper cites.
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira, The non-convex burer-monteiro approach works on smooth semidefinite programs , Advances in Neural Information Processing Systems, 2016, pp. 2757–2765
2016
Earlier work this paper cites.
Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep learning , MIT press, 2016
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
Rong Ge, Jason D Lee, and Tengyu Ma, Matrix completion has no spurious local minimum , Advances in Neural Information Processing Systems, 2016, pp. 2973–2981
2016
Cited alongside, same era.
Kenji Kawaguchi, Deep learning without poor local minima , Advances in neural information processing systems, 2016, pp. 586–594
2016
Cited alongside, same era.
Hamed Karimi, Julie Nutini, and Mark Schmidt, Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition , Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer, 2016, pp. 795–811
2016
Cited alongside, same era.
Murat A Erdogdu, Lester Mackey, and Ohad Shamir, Global non-convex optimization with discretized diffusions , Advances in Neural Information Processing Systems, 2018, pp. 9671–9680
2018
Later among the works it cites.
Andreas Elsener and Sara van de Geer, Sharp oracle inequalities for stationary points of nonconvex penalized m-estimators , IEEE Transactions on Information Theory 65
2018
Later among the works it cites.
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang, Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator , Advances in Neural Information Processing Systems, 2018, pp. 689–699
2018
Later among the works it cites.
Yixin Fang, Jinfeng Xu, and Lei Yang, Online bootstrap confidence intervals for the stochastic gradient descent estimator , The Journal of Machine Learning Research 19
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Brosse, Alain Durmus, Éric Moulines, and Marcelo Pereyra, Sampling from a log-concave distribution with compact support with proximal langevin monte carlo , COLT, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Arnak S Dalalyan, Theoretical guarantees for approximate sampling from smooth and log-concave densities , Journal of the Royal Statistical Society: Series B (Statistical Methodology) 79
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Alain Durmus and Eric Moulines, Nonasymptotic convergence analysis for the unadjusted langevin algorithm , The Annals of Applied Probability 27
2017
Cited alongside, same era.
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein, Automated inference with adaptive batches , Artificial Intelligence and Statistics, 2017, pp. 1504–1513
2017
Cited alongside, same era.
Jianqing Fan and Qiwei Yao, The elements of financial econometrics , Cambridge University Press, 2017
2017
Cited alongside, same era.
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan, How to escape saddle points efficiently , Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 1724–1732
2017
Cited alongside, same era.
2018
Later among the works it cites.
Siyuan Ma, Raef Bassily, and Mikhail Belkin, The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning , International Conference on Machine Learning, 2018, pp. 3325–3334
2018
Later among the works it cites.
Song Mei, Yu Bai, and Andrea Montanari, The landscape of empirical risk for nonconvex losses , The Annals of Statistics 46
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Nilesh Tripuraneni, Mitchell Stern, Chi Jin, Jeffrey Regier, and Michael I Jordan, Stochastic cubic regularization for fast nonconvex optimization , Advances in neural information processing systems, 2018, pp. 2899–2908
2018
Later among the works it cites.
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu, Global convergence of langevin dynamics based algorithms for nonconvex optimization , Advances in Neural Information Processing Systems, 2018, pp. 3122–3133
2018
Later among the works it cites.
Andreas Anastasiou, Krishnakumar Balasubramanian, and Murat A Erdogdu, Normal approximation for stochastic gradient descent via non-asymptotic rates of martingale clt , Conference on Learning Theory, 2019, pp. 115–137
2019
Later among the works it cites.
Yuejie Chi, Yue M Lu, and Yuxin Chen, Nonconvex optimization meets low-rank matrix factorization: An overview , IEEE Transactions on Signal Processing 67
2019
Later among the works it cites.
Changxiao Cai, Gen Li, H Vincent Poor, and Yuxin Chen, Nonconvex low-rank tensor completion from noisy data , Advances in Neural Information Processing Systems, 2019, pp. 1861–1872
2019
Later among the works it cites.
2019
Later among the works it cites.
Aymeric Dieuleveut, Alain Durmus, and Francis Bach, Bridging the gap between constant step size stochastic gradient descent and markov chains , The Annals of Statistics (to appear) (2019)
2019
Later among the works it cites.
Xuechen Li, Yi Wu, Lester Mackey, and Murat A Erdogdu, Stochastic runge-kutta accelerates langevin monte carlo and beyond , Advances in Neural Information Processing Systems, 2019, pp. 7748–7760
2019
Later among the works it cites.
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson, A simple baseline for bayesian uncertainty in deep learning , Advances in Neural Information Processing Systems, 2019, pp. 13132–13143
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Sharan Vaswani, Francis Bach, and Mark Schmidt, Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron , The 22nd International Conference on Artificial Intelligence and Statistics, 2019, pp. 1195–1204
2019
Later among the works it cites.
Dong Xia, Ming Yuan, and Cun-Hui Zhang, Statistically optimal and computationally efficient low rank tensor completion from noisy entries , The Annals of Statistics (to appear) (2019)
2019
Later among the works it cites.
Xi Chen, Jason D Lee, Xin T Tong, Yichen Zhang, et al., Statistical inference for model parameters in stochastic gradient descent , The Annals of Statistics 48
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Ruoqi Shen and Yin Tat Lee, The randomized midpoint method for log-concave sampling , Advances in Neural Information Processing Systems, 2019, pp. 2098–2109
2098
Closest in time.