Fetching the paper…
Reading the bibliography…
It is widely conjectured that the reason that training algorithms for neural networks are successful because all local minima lead to similar performance, for example, see (LeCun et al., 2015, Choromanska et al., 2015, Dauphin et al., 2014).
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Eigenfaces vs. fisherfaces: Recognition using class specific linear projection
P. N. Belhumeur, J. P Hespanha, and D. J. Kriegman · 1997
Earlier work this paper cites.
Sparse pca. extracting multi-scale structure from data
C. Chennubhotla and A. Jepson · 2001
Earlier work this paper cites.
Active appearance models
T. F. Cootes, G. J. Edwards, and C. J. Taylor · 2001
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
I. J Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
H. Sedghi and A. Anandkumar · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. E. Hinton · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Y. Li, J. Yosinski, J. Clune, H. Lipson, and J. Hopcroft · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Cited alongside, same era.
Global optimality in tensor factorization, deep learning, and beyond
B. D Haeffele and R. Vidal · 2015
Cited alongside, same era.
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Learning depth-three neural networks in polynomial time
S. Goel and A. Klivans · 2017
Later among the works it cites.
When is a convolutional filter easy to learn?
S. S. Du, J. D. Lee, and Y. Tian · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L Bartlett, and I. S Dhillon · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Later among the works it cites.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
A. Gautier, Q. N. Nguyen, and M. Hein · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Cited alongside, same era.
Learning relus via gradient descent
M. Soltanolkotabi · 2017
Cited alongside, same era.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Cited alongside, same era.
C. Yun, S. Sra, and A. Jadbabaie · 2017
Later among the works it cites.
The loss surface and expressivity of deep convolutional neural networks
Q. Nguyen and M. Hein · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Later among the works it cites.
Theoretical properties of the global optimizer of two layer neural network
D. Boob and G. Lan · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2017
Later among the works it cites.
Densely connected convolutional networks
G Huang, Zhuang L., Kilian Q. W., and Laurens V. D. M · 2017
Later among the works it cites.