Fetching the paper…
Reading the bibliography…
Stochastic Gradient Langevin Dynamics (SGLD) is a popular variant of Stochastic Gradient Descent, where properly scaled isotropic Gaussian noise is added to an unbiased estimate of the gradient at each iteration.
Laplace’s method revisited: weak convergence of probability measures
C.-R. Hwang · 1980
Earlier work this paper cites.
Mimicking the one-dimensional marginal distributions of processes having an Ito differential
I. Gyöngy · 1986
Earlier work this paper cites.
Diffusion for global optimization in R n {\mathbb R}^{n}
T.-S. Chiang, C.-R. Hwang, and S.-J. Sheu · 1987
Earlier work this paper cites.
Asymptotic Behavior of Dissipative Systems
J. K. Hale · 1988
Earlier work this paper cites.
Recursive stochastic algorithms for global optimization in R d {\mathbb R}^{d}
S. B. Gelfand and S. K. Mitter · 1991
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Dynamical Systems and Numerical Analysis
A. M. Stuart and A. R. Humphries · 1996
Earlier work this paper cites.
Convergence rates for annealing diffusion processes
D. Márquez · 1997
Earlier work this paper cites.
Weak convergence rates for stochastic approximation with application to multiple targets and simulated annealing
M. Pelletier · 1998
Earlier work this paper cites.
A strong approximation theorem for stochastic recursive algorithms
V. S. Borkar and S. K. Mitter · 1999
Earlier work this paper cites.
Statistics of Random Processes I: General Theory
R. S. Liptser and A. N. Shiryaev · 2001
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Topics in Optimal Transportation , volume 58 of Graduate Studies in Mathematics
C. Villani · 2003
Cited alongside, same era.
Transportation cost-information inequalities and applications to random dynamical systems and diffusions
H. Djellout, A. Guillin, and L. Wu · 2004
Cited alongside, same era.
Introductory Lectures on Convex Optimization
Y. Nesterov · 2004
Cited alongside, same era.
Weighted Csiszár–Kullback–Pinsker inequalities and applications to transportation inequalities
F. Bolley and C. Villani · 2005
Cited alongside, same era.
Metastability in reversible diffusion processes II. Precise asymptotics for small eigenvalues
A. Bovier, V. Gayrard, and M. Klein · 2005
Cited alongside, same era.
Stability results in learning theory
A. Rakhlin, S. Mukherjee, and T. Poggio · 2005
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W Teh · 2014
Later among the works it cites.
Optimal transport bounds between the time-marginals of a multidimensional diffusion and its Euler scheme
A. Alfonsi, B. Jourdain, and A. Kohatsu-Higa · 2015
Later among the works it cites.
J.-B. Bardet, N. Gozlan, F. Malrieu, and P.-A. Zitt · 2015
Later among the works it cites.
Escaping the local minima via simulated annealing: Optimization of approximately convex functions
A. Belloni, T. Liang, H. Narayanan, and A. Rakhlin · 2015
Later among the works it cites.
Sampling from a log-concave distribution with Projected Langevin Monte Carlo
S. Bubeck, R. Eldan, and J. Lehec · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Elements of Information Theory
T. M. Cover and J. A. Thomas · 2006
Cited alongside, same era.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
S. Mukherjee, P. Niyogi, T. Poggio, and R. Rifkin · 2006
Cited alongside, same era.
A simple proof of the Poincaré inequality for a large class of probability measures including the log-concave case
D. Bakry, F. Barthe, P. Cattiaux, and A. Guillin · 2008
Cited alongside, same era.
A note on Talagrand’s transportation inequality and logarithmic Sobolev inequality
P. Cattiaux, A. Guillin, and L. Wu · 2010
Cited alongside, same era.
Sparse regression learning by aggregation and Langevin Monte Carlo
A. S. Dalalyan and A. B. Tsybakov · 2012
Cited alongside, same era.
Analysis and Geometry of Markov Diffusion Operators
D. Bakry, I. Gentil, and M. Ledoux · 2014
Cited alongside, same era.
A. Durmus and E. Moulines · 2015
Later among the works it cites.
Escaping from saddle points-online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M Hardt, B Recht, and Y Singer · 2015
Later among the works it cites.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2016
Later among the works it cites.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
A. S. Dalalyan · 2016
Later among the works it cites.
On graduated optimization for stochastic non-convex problems
E. Hazan, K. Levi, and S. Shalev-Shwartz · 2016
Later among the works it cites.
Wasserstein continuity of entropy and outer bounds for interference channels
Y. Polyanskiy and Y. Wu · 2016
Later among the works it cites.