Fetching the paper…
Reading the bibliography…
Minimizing non-convex and high-dimensional objective functions is challenging, especially when training modern deep neural networks.
Optimization by simulated annealing
S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi · 1983
Earlier work this paper cites.
Diffusions for global optimization
S. Geman and C. Hwang · 1986
Earlier work this paper cites.
Simulated annealing: Practice versus theory
L. Ingber · 1993
Earlier work this paper cites.
Escaping free-energy minima
A. Laio and M. Parrinello · 2002
Earlier work this paper cites.
The art of molecular dynamics simulation
D. C. Rapaport · 2004
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Convergence of numerical time-averaging and stationary measures via poisson equations
J. C. Mattingly, A. M. Stuart, and M. V. Tretyakov · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P. Manzagol · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton · 2013
Cited alongside, same era.
Stochastic gradient Hamiltonian Monte Carlo
T. Chen, E. B. Fox, and C. Guestrin · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Cited alongside, same era.
On the convergence of stochastic gradient MCMC algorithms with high-order integrators
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Fei-Fei Li · 2015
Later among the works it cites.
Adding gradient noise improves learning for very deep networks
A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens · 2015
Later among the works it cites.
Deep learning with elastic averaging SGD
S. Zhang, A. Choromanska, and Y. LeCun · 2015
Later among the works it cites.
Bridging the gap between stochastic gradient MCMC and stochastic optimization
C. Chen, D. Carlson, Z. Gan, C. Li, and L. Carin · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Deep learning without poor local minima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Chen, N. Ding, and L. Carin · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G.B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Extended hamiltonian approach to continuous tempering
G. Gobbo and B.J. Leimkuhler · 2015
Cited alongside, same era.
K. Kawaguchi · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. Tang · 2016
Later among the works it cites.
Continuous tempering molecular dynamics: A deterministic approach to simulated tempering
N. Lenner and G. Mathias · 2016
Later among the works it cites.