Fetching the paper…
Reading the bibliography…
Training neural networks is a challenging non-convex optimization problem, and backpropagation or gradient descent can get stuck in spurious local optima.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Back propagation fails to separate where perceptrons succeed
Martin L Brady, Raghu Raghavan, and Joseph Slawny · 1989
Earlier work this paper cites.
multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Nelson Morgan and Hervé Bourlard · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
On the problem of local minima in backpropagation
Marco Gori and Alberto Tesi · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim L Blum and Ronald L Rivest · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns and Umesh Virkumar Vazirani · 1994
Earlier work this paper cites.
Fourier analysis and filtering of a single hidden layer perceptron
Robert J Marks II and Payman Arabshahi · 1994
Earlier work this paper cites.
Neural networks: a systematic introduction
Raúl Rojas · 1996
Earlier work this paper cites.
Successes and failures of backpropagation: A theoretical investigation
P Frasconi, M Gori, and A Tesi · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Hardness results for neural network approximation problems
Peter Bartlett and Shai Ben-David · 1999
Earlier work this paper cites.
Hardness results for general two-layer neural networks
Christian Kuhlmann · 2000
Earlier work this paper cites.
Rank-one approximation to high order tensors
T. Zhang and G. Golub · 2001
Cited alongside, same era.
Training a single sigmoidal neuron is hard
Jiří Šíma · 2002
Cited alongside, same era.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Cited alongside, same era.
Graphical models, exponential families, and variational inference
Martin J Wainwright and Michael I Jordan · 2008
Cited alongside, same era.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Cited alongside, same era.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Cited alongside, same era.
Learning Sparsely Used Overcomplete Dictionaries
A. Agarwal, A. Anandkumar, P. Jain, P. Netrapalli, and R. Tandon · 2014
Later among the works it cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
Random design analysis of ridge regression
Daniel Hsu, Sham M Kakade, and Tong Zhang · 2014
Later among the works it cites.
Score Function Features for Discriminative Learning: Matrix and Tensor Frameworks
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Latent variable graphical model selection via convex optimization
Venkat Chandrasekaran, Pablo Parrilo, Alan S Willsky, et al · 2010
Cited alongside, same era.
On autoencoders and score matching for energy based models
Kevin Swersky, David Buchman, Nando D Freitas, Benjamin M Marlin, et al · 2011
Cited alongside, same era.
What regularized auto-encoders learn from the data generating distribution
Guillaume Alain and Yoshua Bengio · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Cited alongside, same era.
A Tensor Spectral Approach to Learning Mixed Membership Community Models
A. Anandkumar, R. Ge, D. Hsu, and S. M. Kakade · 2013
Cited alongside, same era.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2013
Cited alongside, same era.
Pravesh Kothari and Raghu Meka · 2014
Later among the works it cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Later among the works it cites.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2014
Later among the works it cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Later among the works it cites.
Learning Overcomplete Latent Variable Models through Tensor Methods
Animashree Anandkumar, Rong Ge, and Majid Janzamin · 2015
Closest in time.
The loss surface of multilayer networks
Anna Choromanska, Mikael Henaff, Michaël Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Closest in time.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D. Haeffele and René Vidal · 2015
Closest in time.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Closest in time.
Fast and guaranteed tensor decomposition via sketching
Yining Wang, Hsiao-Yu Tung, Alexander Smola, and Animashree Anandkumar · 2015
Closest in time.
ℓ 1 \ell_{1} -regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D. Lee, and Michael I. Jordan · 2015
Closest in time.