Fetching the paper…
Reading the bibliography…
Stochastic gradient Hamiltonian Monte Carlo (SGHMC) is a variant of stochastic gradient with momentum where a controlled and properly scaled Gaussian noise is added to the stochastic gradients to steer the iterates towards a global minimum.
Brownian motion in a field of force and the diffusion model of chemical reactions
Hendrik Anthony Kramers · 1940
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard · 1970
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C Daniel Gelatt, and Mario P Vecchi · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii E Nesterov · 1983
Earlier work this paper cites.
Nonstationary Markov chains and convergence of the annealing algorithm
Basilis Gidas · 1985
Earlier work this paper cites.
A tutorial survey of theory and applications of simulated annealing
Bruce Hajek · 1985
Earlier work this paper cites.
Diffusion for global optimization in ℝ n \mathbb{R}^{n}
Tzuu-Shuh Chiang, Chii-Ruey Hwang, and Shuenn Jyi Sheu · 1987
Earlier work this paper cites.
Hybrid Monte Carlo
Simon Duane, Anthony D Kennedy, Brian J Pendleton, and Duncan Roweth · 1987
Earlier work this paper cites.
Introduction to Optimization
Boris T Polyak · 1987
Earlier work this paper cites.
Asymptotic Behavior of Dissipative Systems
JK Hale · 1988
Earlier work this paper cites.
Asymptotics of the spectral gap with applications to the theory of simulated annealing
Richard A Holley, Shigeo Kusuoka, and Daniel W Stroock · 1989
Earlier work this paper cites.
Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d}
Saul B Gelfand and Sanjoy K Mitter · 1991
Earlier work this paper cites.
Simulated annealing
Dimitris Bertsimas and John Tsitsiklis · 1993
Earlier work this paper cites.
A strong approximation theorem for stochastic recursive algorithms
Vivek S Borkar and Sanjoy K Mitter · 1999
Earlier work this paper cites.
Choosing regularization parameters in iterative methods for ill-posed problems
Misha E Kilmer and Dianne P O’Leary · 2001
Earlier work this paper cites.
Ergodicity for SDEs and approximations: locally Lipschitz vector fields and degenerate noise
Jonathan C Mattingly, Andrew M Stuart, and Desmond J Higham · 2002
Earlier work this paper cites.
Stochastic Differential Equations: An Introduction with Applications
B. K. Øksendal · 2003
Earlier work this paper cites.
Isotropic hypoellipticity and trend to equilibrium for the Fokker-Planck equation with a high-degree potential
Frédéric Hérau and Francis Nier · 2004
Earlier work this paper cites.
Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues
Anton Bovier, Véronique Gayrard, and Markus Klein · 2005
Earlier work this paper cites.
Weighted Csiszár-Kullback-pinsker inequalities and applications to transportation inequalities
François Bolley and Cédric Villani · 2005
Earlier work this paper cites.
Trading convexity for scalability
Ronan Collobert, Fabian Sinz, Jason Weston, and Léon Bottou · 2006
Earlier work this paper cites.
Robust truncated hinge loss support vector machines
Yichao Wu and Yufeng Liu · 2007
Earlier work this paper cites.
Optimal Transport: Old and New
Cédric Villani · 2008
Earlier work this paper cites.
Tighter bounds for structured estimation
Olivier Chapelle, Chuong B Do, Choon H Teo, Quoc V Le, and Alex J Smola · 2009
Earlier work this paper cites.
MCMC using Hamiltonian dynamics. Handbook of Markov Chain Monte Carlo (S. Brooks, A. Gelman, G. Jones, and X.-L. Meng, eds.), 2010
RM Neal · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian posterior sampling via stochastic gradient Fisher scoring
Sungjin Ahn, Anoop Korattikara, and Max Welling · 2012
Earlier work this paper cites.
Robust greedy algorithms for compressed sensing
S Alireza Razavi, Esa Ollila, and Visa Koivunen · 2012
Earlier work this paper cites.
Sonja Cox, Martin Hutzenthaler, and Arnulf Jentzen · 2013
Earlier work this paper cites.
An Introduction to Statistical Learning
Gareth James, Daniela Witten, Trevor Hastie, and Robert Tibshirani · 2013
Earlier work this paper cites.
Statistics of Random Processes: I. General Theory
Robert S Liptser and Albert N Shiryaev · 2013
Cited alongside, same era.
Algorithms for direct 0–1 loss optimization in binary classification
Tan Nguyen and Scott Sanner · 2013
Cited alongside, same era.
Stochastic gradient Riemannian Langevin dynamics on the probability simplex
Sam Patterson and Yee Whye Teh · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Robust variable selection with exponential squared loss
Xueqin Wang, Yunlu Jiang, Mian Huang, and Heping Zhang · 2013
Cited alongside, same era.
Optimizing the integrator step size for Hamiltonian Monte Carlo
Rapid Mixing of Hamiltonian Monte Carlo on Strongly Log-Concave Distributions
O. Mangoubi and A. Smith · 2017
Later among the works it cites.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient Langevin dynamics
Yuchen Zhang, Percy. Liang, and Moses. Charikar · 2017
Later among the works it cites.
A nonconvex approach for phase retrieval: Reshaped Wirtinger flow and incremental algorithms
Huishuai Zhang, Yi Zhou, Yingbin Liang, and Yuejie Chi · 2017
Later among the works it cites.
Sharp Convergence Rates for Langevin Dynamics in the Nonconvex Setting
X. Cheng, N. S. Chatterji, Y. Abbasi-Yadkori, P. L. Bartlett, and M. I. Jordan · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MJ Betancourt, Simon Byrne, and Mark Girolami · 2014
Cited alongside, same era.
Stochastic gradient Hamiltonian Monte Carlo
Tianqi Chen, Emily Fox, and Carlos Guestrin · 2014
Cited alongside, same era.
Bayesian sampling using stochastic gradient thermostats
Nan Ding, Youhan Fang, Ryan Babbush, Changyou Chen, Robert D Skeel, and Hartmut Neven · 2014
Cited alongside, same era.
Stochastic Processes and Applications: Diffusion Processes, the Fokker-Planck and Langevin Equations
Grigorios A Pavliotis · 2014
Cited alongside, same era.
A differential equation for modeling Nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
Escaping the local minima via simulated annealing: Optimization of approximately convex functions
Alexandre Belloni, Tengyuan Liang, Hariharan Narayanan, and Alexander Rakhlin · 2015
Cited alongside, same era.
On the convergence of stochastic gradient MCMC algorithms with high-order integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Cited alongside, same era.
Underdamped Langevin MCMC: A non-asymptotic analysis
Xiang Cheng, Niladri S Chatterji, Peter L Bartlett, and Michael I Jordan · 2018
Closest in time.
Accelerated methods for nonconvex optimization
Y. Carmon, J. Duchi, O. Hinder, and A. Sidford · 2018
Closest in time.
On the theory of variance reduction for stochastic gradient Monte Carlo
Niladri S Chatterji, Nicolas Flammarion, Yi-An Ma, Peter L Bartlett, and Michael I Jordan · 2018
Closest in time.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
Simon S Du, Jason D Lee, Yuandong Tian, Aarti Singh, and Barnabas Poczos · 2018
Closest in time.
Uniform convergence of gradients for non-convex learning and optimization
Dylan J Foster, Ayush Sekhari, and Karthik Sridharan · 2018
Closest in time.
Tianyi Liu, Zhehui Chen, Enlu Zhou, and Tuo Zhao · 2018
Closest in time.
Accelerating greedy coordinate descent methods
Haihao Lu, Robert M Freund, and Vahab Mirrokni · 2018
Closest in time.
Convergence rate of Riemannian Hamiltonian Monte Carlo and faster polytope volume computation
Yin Tat Lee and Santosh S Vempala · 2018
Closest in time.
The landscape of empirical risk for nonconvex losses
Song Mei, Yu Bai, and Andrea Montanari · 2018
Closest in time.
Does Hamiltonian Monte Carlo mix faster than a random walk on multimodal densities?
Oren Mangoubi, S. Pillai, Natesh, and Aaron Smith · 2018
Closest in time.
Understanding the acceleration phenomenon via high-resolution differential equations
Bin Shi, Simon S Du, Michael I Jordan, and Weijie J Su · 2018
Closest in time.
Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization
U. Şimşekli, Ç. Yıldız, T. H. Nguyen, G. Richard, and A. Taylan Cemgil · 2018
Closest in time.
Local optimality and generalization guarantees for the Langevin algorithm via empirical metastability
Belinda Tzen, Tengyuan Liang, and Maxim Raginsky · 2018
Closest in time.
Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem
Andre Wibisono · 2018
Closest in time.
Global convergence of Langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Closest in time.
Sorting out Lipschitz function approximation
Cem Anil, James Lucas, and Roger Grosse · 2019
Closest in time.
User-friendly guarantees for the Langevin Monte Carlo with inaccurate gradient
Arnak S. Dalalyan and Avetik G. Karagulyan · 2019
Closest in time.
Bounding the error of discretized Langevin algorithms for non-strongly log-concave targets
Arnak S Dalalyan, Avetik Karagulyan, and Lionel Riou-Durand · 2019
Closest in time.
Hypoelliptic diffusions: filtering and inference from complete and partial observations
Susanne Ditlevsen and Adeline Samson · 2019
Closest in time.
Couplings and quantitative contraction rates for Langevin dynamics
Andreas Eberle, Arnaud Guillin, and Raphael Zimmer · 2019
Closest in time.
On variance reduction for stochastic smooth convex optimization with multiplicative noise
A. Jofré and P. Thompson · 2019
Closest in time.
Behavior of accelerated gradient methods near critical points of nonconvex problems
Michael O’Neill and Stephen J Wright · 2019
Closest in time.
Differentially private empirical risk minimization with non-convex loss functions
Di Wang, Changyou Chen, and Jinhui Xu · 2019
Closest in time.
On sampling from a log-concave density using kinetic Langevin diffusions
Arnak S Dalalyan and Lionel Riou-Durand · 2020
Closest in time.
Breaking reversibility accelerates Langevin dynamics for global non-convex optimization
Xuefeng Gao, Mert Gürbüzbalaban, and Lingjiong Zhu · 2020
Closest in time.