Fetching the paper…
Reading the bibliography…
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. Polyak and A. Juditsky · 1992
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
M. Marcus, M.-A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Fast exact multiplication by the hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Weight space structure and internal representations: A direct approach to learning and generalization in multilayer neural networks
R. Monasson and R. Zecchina · 1995
Earlier work this paper cites.
Analytical and numerical study of internal representations in multilayer neural networks with binary weights
S. Cocco, R. Monasson, and R. Zecchina · 1996
Earlier work this paper cites.
Mutual information, metric entropy and cumulative relative entropy risk
D. Haussler and M. Opper · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Langevin diffusions and Metropolis-Hastings algorithms
G. Roberts and O. Stramer · 2002
Earlier work this paper cites.
Metastability: A potential theoretic approach
A. Bovier and F. den Hollander · 2006
Earlier work this paper cites.
The statistics of critical points of Gaussian fields on large-dimensional spaces
A. Bray and D. Dean · 2007
Earlier work this paper cites.
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Y. Fyodorov and I. Williams · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
R. Neal · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Numerical continuation methods: an introduction , volume 13
E. L. Allgower and K. Georg · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Lecture 6.5: RmsProp, Coursera: Neural networks for machine learning
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Maxout networks
I. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
All of statistics: A concise course in statistical inference
L. Wasserman · 2013
Cited alongside, same era.
Stochastic Gradient Hamiltonian Monte Carlo
T. Chen, E. Fox, and C. Guestrin · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
Bayesian sampling using stochastic gradient thermostats
N. Ding, Y. Fang, R. Babbush, C. Chen, R. D. Skeel, and H. Neven · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Later among the works it cites.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and L. Fei-Fei · 2015
Later among the works it cites.
A complete recipe for stochastic gradient MCMC
Y.-A. Ma, T. Chen, and E. Fox · 2015
Later among the works it cites.
On the link between Gaussian homotopy continuation and convex envelopes
H. Mobahi and J. Fisher III · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. Saxe, J. McClelland, and S. Ganguli · 2014
Cited alongside, same era.
Striving for simplicity: The all convolutional net
J. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
Bayesian dark knowledge
A. Balan, V. Rathod, K. Murphy, and M. Welling · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Later among the works it cites.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Closest in time.
Local entropy as a measure for sampling solutions in constraint satisfaction problems
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2016
Closest in time.
Variational inference: A review for statisticians
D. Blei, A. Kucukelbir, and J. McAuliffe · 2016
Closest in time.
T. Cooijmans, N. Ballas, C. Laurent, and A. Courville · 2016
Closest in time.
Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
Z. Gan, C. Li, C. Chen, Y. Pu, Q. Su, and L. Carin · 2016
Closest in time.
C. Gulcehre, M. Moczulski, F. Visin, and Y. Bengio · 2016
Closest in time.
On graduated optimization for stochastic non-convex problems
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2016
Closest in time.
Deep learning without poor local minima
K. Kawaguchi · 2016
Closest in time.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Closest in time.
Training Recurrent Neural Networks by Diffusion
H. Mobahi · 2016
Closest in time.
Singularity of the Hessian in Deep Learning
L. Sagun, L. Bottou, and Y. LeCun · 2016
Closest in time.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. Kingma · 2016
Closest in time.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Closest in time.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Closest in time.