Fetching the paper…
Reading the bibliography…
Gradient descent (GD) is known to converge quickly for convex objective functions, but it can be trapped at local minima.
Optimization by simulated annealing
Scott Kirkpatrick, C Daniel Gelatt, and Mario P Vecchi · 1983
Earlier work this paper cites.
Nonstationary markov chains and convergence of the annealing algorithm
B. Gidas · 1985
Earlier work this paper cites.
Replica Monte Carlo simulation of spin-glasses
Robert H Swendsen and Jian-Sheng Wang · 1986
Earlier work this paper cites.
Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d}
S.B. Gelfand and S.K · 1991
Earlier work this paper cites.
Simulated tempering: a new Monte Carlo scheme
Enzo Marinari and Giorgio Parisi · 1992
Earlier work this paper cites.
Simulated annealing: A proof of convergence
Vincent Granville, Mirko Krivánek, and J-P Rasson · 1994
Earlier work this paper cites.
Annealing Markov Chain Monte Carlo with applications to ancestral inference
Charles J. Geyer and Elizabeth A. Thompson · 1995
Earlier work this paper cites.
Weak convergence rates for stochastic approximation with application to multiple targets and simulated annealing
M. Pelletier · 1998
Earlier work this paper cites.
A parameter study for differential evolution
Roger Gämperle, Sibylle D Müller, and Petros Koumoutsakos · 2002
Earlier work this paper cites.
On the use of particle swarm optimization with multimodal functions
Susana C Esquivel and CA Coello Coello · 2003
Earlier work this paper cites.
Parallel tempering: Theory, applications, and new perspectives
D. J. Earl and M. W. Deem · 2005
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. Nesterov · 2005
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Conditions for rapid mixing of parallel and simulated termpering on multimodal distribution
D. B. Woodard, S. C. Schmidler, and M. Huber · 2009
Earlier work this paper cites.
On the infinite swapping limit for parallel tempering
P. Dupuis, Y. Liu, N. Plattner, and J.D. Doll · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2013
Earlier work this paper cites.
A competitive swarm optimizer for large scale optimization
Ran Cheng and Yaochu Jin · 2014
Earlier work this paper cites.
Likelihood-informed dimension reduction for nonlinear inverse problems
Tiangang Cui, James Martin, Youssef M Marzouk, Antti Solonen, and Alessio Spantini · 2014
Cited alongside, same era.
A trust region algorithm with a worst-case iteration complexity of O( ϵ − 3 / 2 \epsilon^{-3/2} ) for nonconvex optimization
Frank E Curtis, Daniel P Robinson, and Mohammadreza Samadi · 2014
Cited alongside, same era.
Spectral gaps for a Metropolis–Hastings algorithm in infinite dimensions
M. Hairer, A.M. Stuart, and S.J. Vollmer · 2014
Cited alongside, same era.
Poincaré and logarithmic sobolev inequalities by decomposition of the energy landscape
Georg Menz and André Schlichting · 2014
Cited alongside, same era.
A complete recipe for stochastic gradient MCMC
Yi-An Ma, Tianqi Chen, and Emily B. Fox · 2015
Cited alongside, same era.
Global optimality of local search for low rank matrix recovery
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Later among the works it cites.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
Simon Du, Jason Lee, Yuandong Tian, Aarti Singh, and Barnabas Poczos · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2018
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I Jordan · 2018
Later among the works it cites.
Beyond log-concavity: Provable guarantees for sampling multi-modal distributions using simulated tempering langevin monte carlo
Holden Lee, Andrej Risteski, and Rong Ge · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Cited alongside, same era.
Theoretical guarantees for approximate sampling from a smooth and log-concave density
A.S. Dalalyan · 2017
Cited alongside, same era.
Nonasymptotic convergence analysis for the unadjusted Langevin algorithm
A. Durmus and E. Moulines · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Cited alongside, same era.
Solving SDPs for synchronization and maxcut problems via the Grothendieck inequality
Song Mei, Theodor Misiakiewicz, Andrea Montanari, and Roberto I Oliveira · 2017
Cited alongside, same era.
Non-square matrix sensing without spurious local minima via the Burer-Monteiro approach
Dohyung Park, Anastasios Kyrillidis, Constantine Carmanis, and Sujay Sanghavi · 2017
Cited alongside, same era.
Oren Mangoubi and Nisheeth K Vishnoi · 2018
Later among the works it cites.
Global convergence of Langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Later among the works it cites.
Third-order smoothness helps: Faster stochastic optimization algorithms for finding local minima
Yaodong Yu, Pan Xu, and Quanquan Gu · 2018
Later among the works it cites.
On stationary-point hitting time and ergodicity of stochastic gradient Langevin dynamics
X. Chen, S. Du, and X. T. Tong · 2019
Later among the works it cites.
Accelerating nonconvex learning via replica exchange Langevin diffusion
Y. Chen, J. Chen, J. Dong, J. Peng, and Z. Wang · 2019
Later among the works it cites.
User-friendly guarantees for the Langevin Monte Carlo with inaccurate gradient
A.S. Dalalyan and A. Karagulyan · 2019
Later among the works it cites.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2019
Later among the works it cites.
Sampling can be faster than optimization
Y. Ma, Y. Chen, C. Jin, N. Flammarion, and M. I. Jordan · 2019
Later among the works it cites.
Localization for mcmc: sampling high-dimensional posterior distributions with local structure
Matthias Morzfeld, Xin T Tong, and Youssef M Marzouk · 2019
Later among the works it cites.
Rapid convergence of the unadjusted langevin algorithm: Isoperimetry suffices
Santosh S Vempala and Andre Wibisono · 2019
Later among the works it cites.
Spectral gap of replica exchange langevin diffusion on mixture distributions
Jing Dong and Xin T Tong · 2020
Closest in time.
Mala-within-gibbs samplers for high-dimensional distributions with sparse conditional structure
Xin T. Tong, Mathias Morzfeld, and Youssef M. Marzouk · 2020
Closest in time.