Fetching the paper…
Reading the bibliography…
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Neural Networks for Pattern Recognition
C. M. Bishop · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
D. Saad and S. A. Solla · 1995
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Earlier work this paper cites.
Statistical mechanics of learning
A. Engel and C. Van den Broeck · 2001
Earlier work this paper cites.
Gaussian processes in machine learning
C. E. Rasmussen · 2003
Earlier work this paper cites.
On-line learning theory of soft committee machines with correlated hidden units - a steepest gradient descent and natural gradient descent
M. Inoue, H. Park, and M. Okada · 2003
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
Explorations on high dimensional landscapes
L. Sagun, V. U. Guney, G. B. Arous, and Y. LeCun · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Deep learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
L. Sagun, L. Bottou, and Y. LeCun · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C. D. Freeman and J. Bruna · 2016
Cited alongside, same era.
Depth creates no bad local minima
H. Lu and K. Kawaguchi · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Later among the works it cites.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Cited alongside, same era.
Energy landscapes for machine learning
A. J. Ballard, R. Das, S. Martiniani, D. Mehta, L. Sagun, J. D. Stevenson, and D. J. Wales · 2017
Cited alongside, same era.
F. Draxler, K. Veschgini, M. Salmhofer, and F. A. Hamprecht · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
S. Spigler, M. Geiger, S. d’Ascoli, L. Sagun, G. Biroli, and M. Wyart · 2018
Later among the works it cites.
Comparing dynamics: Deep neural networks versus glassy systems
M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli · 2018
Later among the works it cites.