Fetching the paper…
Reading the bibliography…
Replica exchange stochastic gradient Langevin dynamics (reSGLD) has shown promise in accelerating the convergence in non-convex learning; however, an excessively large correction for avoiding biases from noisy energy estimators has limited the potential of the acceleration.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The Theory of Stochastic Processes I
Iosif I. Gikhman and Anatoli V. Skorokhod · 1980
Earlier work this paper cites.
Mimicking the One-dimensional Marginal Distributions of Processes Having an Itô differential
István Gyöngy · 1986
Earlier work this paper cites.
Replica Monte Carlo Simulation of Spin-Glasses
Robert H. Swendsen and Jian-Sheng Wang · 1986
Earlier work this paper cites.
On the Dirichlet Problem for a Class of Second Order PDE Systems with Small Parameter
A. Eizenberg and M. Freidlin · 1990
Earlier work this paper cites.
Ergodicity for SDEs and Approximations: Locally Lipschitz Vector Fields and Degenerate Noise
J.C. Mattingly, A.M. Stuartb, and D.J. Highamc · 2002
Earlier work this paper cites.
Stochastic Differential Equations: An Introduction with Applications
B. Øksendal · 2003
Earlier work this paper cites.
Weighted Csiszár-Kullback-Pinsker Inequalities and Applications to Transportation Inequalities
François Bolley and Cédric Villani · 2005
Earlier work this paper cites.
Parallel Tempering: Theory, Applications, and New Perspectives
David J. Earl and Michael W. Deem · 2005
Earlier work this paper cites.
The Pseudo-Marginal Approach for Efficient Monte Carlo Computations
Christophe Andrieu and Gareth O. Roberts · 2009
Earlier work this paper cites.
Mimicking the Marginal Distributions of a Semimartingale
Amel Bentata and Rama Cont · 2009
Earlier work this paper cites.
Hybrid Switching Diffusions: Properties and Applications
George Yin and Chao Zhu · 2010
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Analysis and Geometry of Markov Diffusion Operators
Dominique Bakry, Ivan Gentil, and Michel Ledoux · 2014
Cited alongside, same era.
Stochastic Gradient Hamiltonian Monte Carlo
Tianqi Chen, Emily B. Fox, and Carlos Guestrin · 2014
Cited alongside, same era.
SAGA: A Fast Incremental Gradient Method with Support for Non-Strongly Convex Composite Objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
On the Convergence of Stochastic Gradient MCMC Algorithms with High-order Integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Cited alongside, same era.
Stop Wasting My Gradients: Practical SVRG
Reza Harikandeh, Mohamed Osama Ahmed, Alim Virani, Mark Schmidt, Jakub Konečný, and Scott Sallinen · 2015
Cited alongside, same era.
Variance Reduction in Stochastic Gradient Langevin Dynamics
Avinava Dubey, Sashank J. Reddi, Barnabás Póczos, Alexander J. Smola, Eric P. Xing, and Sinead A. Williamson · 2016
A Hitting Time Analysis of Stochastic Gradient Langevin Dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Later among the works it cites.
On the Theory of Variance Reduction for Stochastic Gradient Monte Carlo
Niladri Chatterji, Nicolas Flammarion, Yi-An Ma, Peter Bartlett, and Michael Jordan · 2018
Later among the works it cites.
Beyond Log-concavity: Provable Guarantees for Sampling Multi-modal Distributions using Simulated Tempering Langevin Monte Carlo
Holden Lee, Andrej Risteski, and Rong Ge · 2018
Later among the works it cites.
Visualizing the Loss Landscape of Neural Nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
Global Convergence of Langevin Dynamics Based Algorithms for Nonconvex Optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Later among the works it cites.
Control Variates for Stochastic Gradient MCMC
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
Chunyuan Li, Changyou Chen, David Carlson, and Lawrence Carin · 2016
Cited alongside, same era.
Consistency and Fluctuations for Stochastic Gradient Langevin Dynamics
Yee Whye Teh, Alexandre Thiery, and Sebastian Vollmer · 2016
Cited alongside, same era.
Exploration of the (Non-) Asymptotic Bias and Variance of Stochastic Gradient Langevin Dynamics
Sebastian J. Vollmer, Konstantinos C. Zygalakis, and Yee Whye Teh · 2016
Cited alongside, same era.
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensemble
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Jack Baker, Paul Fearnhead, Emily B. Fox, and Christopher Nemeth · 2019
Later among the works it cites.
Accelerating Nonconvex Learning via Replica Exchange Langevin Diffusion
Yi Chen, Jinglin Chen, Jing Dong, Jian Peng, and Zhaoran Wang · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Speeding Up MCMC by Efficient Data Subsampling
Matias Quiroz, Robert Kohn, Mattias Villani, and Minh-Ngoc Tran · 2019
Later among the works it cites.
Stochastic Nested Variance Reduction for Nonconvex Optimization
Dongruo Zhou, Pan Xu, and Quanquan Gu · 2019
Later among the works it cites.
Non-Convex Learning via Replica Exchange Stochastic Gradient MCMC
Wei Deng, Qi Feng, Liyao Gao, Faming Liang, and Guang Lin · 2020
Closest in time.
Spectral Gap of Replica Exchange Langevin Diffusion on Mixture Distributions
Jing Dong and Xin T. Tong · 2020
Closest in time.
Accelerating the Diffusion-based Ensemble Sampling by Non-reversible Dynamics
Futoshi Futami, Issei Sato, and Masashi Sugiyama · 2020
Closest in time.
Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
Ruqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen, and Andrew Gordon Wilson · 2020
Closest in time.