Fetching the paper…
Reading the bibliography…
We study the complexity of training neural network models with one hidden nonlinear activation layer and an output weighted sum layer.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L. Rivest · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael J. Kearns · 1993
Earlier work this paper cites.
Geometric applications of Fourier series and spherical harmonics
Helmut Groemer · 1996
Earlier work this paper cites.
On learning correlated boolean functions using statistical queries
Ke Yang · 2001
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Earlier work this paper cites.
Statistical algorithms and a lower bound for planted clique
Vitaly Feldman, Elena Grigorescu, Lev Reyzin, Santosh Vempala, and Ying Xiao · 2013
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
Generalization bounds for neural networks through tensor factorization
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
Complexity theoretic limitations on learning DNF’s
Amit Daniely and Shai Shalev-Shwartz · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2016
Cited alongside, same era.
Cryptographic hardness of learning
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D. Lee, and Tengyu Ma · 2017
Later among the works it cites.
Reliably Learning the ReLU in Polynomial Time
Surbhi Goel, Varun Kanade, Adam R. Klivans, and Justin Thaler · 2017
Later among the works it cites.
Learning depth-three neural networks in polynomial time
Surbhi Goel and Adam R. Klivans · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
On the complexity of learning neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam R. Klivans · 2016
Cited alongside, same era.
Provable tensor methods for learning mixtures of generalized linear models
Hanie Sedghi, Majid Janzamin, and Anima Anandkumar · 2016
Cited alongside, same era.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Le Song, Santosh Vempala, John Wilmes, and Bo Xie · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett, and Inderjit S. Dhillon · 2017
Later among the works it cites.
The total variation distance between high-dimensional gaussians
Luc Devroye, Abbas Mehrabian, and Tommy Reddad · 2018
Closest in time.
Learning two-layer neural networks with symmetric inputs
Rong Ge, Rohith Kuditipudi, Zhize Li, and Xiang Wang · 2018
Closest in time.
Learning one convolutional layer with overlapping patches
Surbhi Goel, Adam Klivans, and Raghu Meka · 2018
Closest in time.
On the spectral bias of deep neural networks
Nasim Rahaman, Devansh Arpit, Aristide Baratin, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2018
Closest in time.