Fetching the paper…
Reading the bibliography…
We study the Stochastic Gradient Langevin Dynamics (SGLD) algorithm for non-convex optimization.
A lower bound for the smallest eigenvalue of the laplacian
J. Cheeger · 1969
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
A. L. Blum and R. L. Rivest · 1992
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
D. Haussler · 1992
Earlier work this paper cites.
The hardness of approximate optima in lattices, codes, and systems of linear equations
S. Arora, L. Babai, J. Stern, and Z. Sweedyk · 1993
Earlier work this paper cites.
Random walks in a convex body and an improved volume algorithm
L. Lovász and M. Simonovits · 1993
Earlier work this paper cites.
Exponential convergence of langevin distributions and their discrete approximations
G. O. Roberts and R. L. Tweedie · 1996
Earlier work this paper cites.
An overview of statistical learning theory
V. N. Vapnik · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
B. Laurent and P. Massart · 2000
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2003
Earlier work this paper cites.
Metastability in reversible diffusion processes i: Sharp asymptotics for capacities and exit times
A. Bovier, M. Eckhoff, V. Gayrard, and M. Klein · 2004
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Cited alongside, same era.
The multivariate normal distribution
Y. L. Tong · 2012
Cited alongside, same era.
Advances in optimizing recurrent networks
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Optimal computational and statistical rates of convergence for sparse nonconvex learning problems
Z. Wang, H. Liu, and T. Zhang · 2014
Cited alongside, same era.
Efficient learning of linear separators under bounded noise
P. Awasthi, M. Balcan, N. Haghtalab, and R. Urner · 2015
Cited alongside, same era.
Regularized m-estimators with nonconvexity: statistical and algorithmic theory for local optima
P.-L. Loh and M. J. Wainwright · 2015
Later among the works it cites.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2015
Later among the works it cites.
Finding local minima for nonconvex optimization in linear time
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2016
Later among the works it cites.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Later among the works it cites.
Guarantees in wasserstein distance for the langevin monte carlo algorithm
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finite-time analysis of projected langevin monte carlo
S. Bubeck, R. Eldan, and J. Lehec · 2015
Cited alongside, same era.
Optimal rates for zero-order convex optimization: the power of two function evaluations
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and A. Wibisono · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
Ł. Kaiser and I. Sutskever · 2015
Cited alongside, same era.
K. Kurach, M. Andrychowicz, and I. Sutskever · 2015
Cited alongside, same era.
Neural programmer: Inducing latent programs with gradient descent
A. Neelakantan, Q. V. Le, and I. Sutskever
Cited in the paper.
T. Bonis · 2016
Later among the works it cites.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
A. S. Dalalyan · 2016
Later among the works it cites.
On local maxima in the population likelihood of gaussian mixture models: Structural results and algorithmic consequences
C. Jin, Y. Zhang, S. Balakrishnan, M. J. Wainwright, and M. I. Jordan · 2016
Later among the works it cites.
Gradient descent converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Later among the works it cites.
Consistency and fluctuations for stochastic gradient langevin dynamics
Y. W. Teh, A. H. Thiery, and S. J. Vollmer · 2016
Later among the works it cites.
A comprehensive study of deep bidirectional lstm rnns for acoustic modeling in speech recognition
A. Zeyer, P. Doetsch, P. Voigtlaender, R. Schlüter, and H. Ney · 2016
Later among the works it cites.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Closest in time.