Fetching the paper…
Reading the bibliography…
One of the main difficulties in analyzing neural networks is the non-convexity of the loss function which may have many bad local minima.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Support-vector networks
C. Cortes and V. Vapnik · 1995
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
The best rank-1 approximation of a symmetric tensor and related spherical optimization problems
X. Zhang, C. Ling, and L. Qi · 2012
Earlier work this paper cites.
I. J Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
H. Sedghi and A. Anandkumar · 2014
Earlier work this paper cites.
Structured low-rank matrix factorization: Optimality, algorithm, and applications to image processing
B. Haeffele, E. Young, and R. Vidal · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. E. Hinton · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
B. D Haeffele and R. Vidal · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Cited alongside, same era.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
A. Gautier, Q. N. Nguyen, and M. Hein · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Later among the works it cites.
Learning relus via gradient descent
M. Soltanolkotabi · 2017
Later among the works it cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Later among the works it cites.
Learning depth-three neural networks in polynomial time
S. Goel and A. Klivans · 2017
Later among the works it cites.
When is a convolutional filter easy to learn?
S. S. Du, J. D. Lee, and Y. Tian · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Densely connected convolutional networks
G. Huang and Z. Liu · 2017
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2017
Cited alongside, same era.
The loss surface and expressivity of deep convolutional neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
The loss surface and expressivity of deep convolutional neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L Bartlett, and I. S Dhillon · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
S. S Du and J. D Lee · 2018
Closest in time.
Learning one-hidden-layer neural networks with landscape design
R. Ge, J. D Lee, and T. Ma · 2018
Closest in time.
Understanding the loss surface of neural networks for binary classification
S. Liang, R. Sun, Y. Li, and R. Srikant · 2018
Closest in time.
Are resnets provably better than linear predictors?
O. Shamir · 2018
Closest in time.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.