Fetching the paper…
Reading the bibliography…
Stochastic gradient Langevin dynamics (SGLD) is a fundamental algorithm in stochastic optimization.
Markov chains and stochastic stability
S. Meyn and R. Tweedie · 1993
Earlier work this paper cites.
Ergodicity for sdes and approximations: Locally lipschitz vector fields and degenerate noise
J. C. Mattingly, A. M. Stuart, and D. J. Higham · 2002
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Inverse problems: a bayesian perspective
A. Stuart · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian posterior sampling via stochastic gradient Fisher scoring
Sungjin Ahn, Anoop Korattikara, and Max Welling · 2012
Earlier work this paper cites.
A trust region algorithm with a worst-case iteration complexity of O( ϵ − 3 / 2 \epsilon^{-3/2} ) for nonconvex optimization
Frank E Curtis, Daniel P Robinson, and Mohammadreza Samadi · 2014
Earlier work this paper cites.
On the convergence of stochastic gradient MCMC algorithms with high-order integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Neural gpus learn algorithms
Łukasz Kaiser and Ilya Sutskever · 2015
Earlier work this paper cites.
A complete recipe for stochastic gradient MCMC
Yi-An Ma, Tianqi Chen, and Emily Fox · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Earlier work this paper cites.
Gradient descent efficiently finds the cubic-regularized non-convex newton step
Yair Carmon and John C Duchi · 2016
Earlier work this paper cites.
Variance reduction in stochastic gradient Langevin dynamics
Kumar Avinava Dubey, Sashank J Reddi, Sinead A Williamson, Barnabas Poczos, Alexander J Smola, and Eric P Xing · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Neural random-access machines
Karol Kurach, Marcin Andrychowicz, and Ilya Sutskever · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Neural programmer: Inducing latent programs with gradient descent
Arvind Neelakantan, Quoc V Le, and Ilya Sutskever · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Cited alongside, same era.
A hitting time analysis of stochastic gradient Langevin dyanmics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Later among the works it cites.
Natasha 2: Faster non-convex optimization than SGD
Zeyuan Allen-Zhu · 2018
Later among the works it cites.
Neon2: Finding local minima via first-order oracles
Zeyuan Allen-Zhu and Yuanzhi Li · 2018
Later among the works it cites.
Sampling from a log-concave distribution with projected Langevin Monte Carlo
Sébastien Bubeck, Ronen Eldan, and Joseph Lehec · 2018
Later among the works it cites.
Accelerated methods for nonconvex optimization
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2018
Later among the works it cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Cited alongside, same era.
Sampling from a log-concave distribution with compact support with proximal Langevin Monte Carlo
Nicolas Brosse, Alain Durmus, Éric Moulines, and Marcelo Pereyra · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Aarti Singh, and Barnabas Poczos · 2017
Cited alongside, same era.
Nonasymptotic convergence analysis for the unadjusted Langevin algorithm
Alain Durmus, Eric Moulines, et al · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Simon Du and Jason Lee · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2018
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I Jordan · 2018
Later among the works it cites.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2018
Later among the works it cites.
Local optimality and generalization guarantees for the Langevin algorithm via empirical metastability
Belinda Tzen, Tengyuan Liang, and Maxim Raginsky · 2018
Later among the works it cites.
Third-order smoothness helps: Faster stochastic optimization algorithms for finding local minima
Yaodong Yu, Pan Xu, and Quanquan Gu · 2018
Later among the works it cites.
A critical view of global optimality in deep learning
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
User-friendly guarantees for the Langevin Monte Carlo with inaccurate gradient
Arnak S Dalalyan and Avetik Karagulyan · 2019
Closest in time.
Sharp analysis for nonconvex sgd escaping from saddle points
Cong Fang, Zhouchen Lin, and Tong Zhang · 2019
Closest in time.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2019
Closest in time.
Replica exchange for non-convex optimization
J. Dong and X. T. Tong · 2020
Closest in time.