Approximation by entire functions
Wilfred Kaplan · 1955
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
On the problem of local minima in backpropagation
Marco Gori and Alberto Tesi · 1992
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
Original
H. Sedghi and A. Anandkumar · 2014
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Original
Matus Telgarsky · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Original
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Original
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Original
B. D Haeffele and R. Vidal · 2015
Earlier work this paper cites.
The zero set of a real analytic function
Original
Boris Mityagin · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
A. Gautier, Q. N. Nguyen, and M. Hein · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
Gaussian error linear units (GeLUs)
Original
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Original
A. Brutzkus and A. Globerson · 2017
Earlier work this paper cites.
Learning ReLUs via gradient descent
M. Soltanolkotabi · 2017
Earlier work this paper cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Original
D. Soudry and E. Hoffer · 2017
Earlier work this paper cites.
Learning depth-three neural networks in polynomial time
Original
S. Goel and A. Klivans · 2017
Earlier work this paper cites.
Theoretical properties of the global optimizer of two layer neural network
Original
D. Boob and G. Lan · 2017
Earlier work this paper cites.
When is a convolutional filter easy to learn?
Original
S. S. Du, J. D. Lee, and Y. Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L Bartlett, and I. S Dhillon · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with ReLU activation
Y. Li and Y. Yuan · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Earlier work this paper cites.
Global optimality conditions for deep neural networks
Original
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2017
Earlier work this paper cites.
The loss surface and expressivity of deep convolutional neural networks
Original
Q. Nguyen and M. Hein · 2017
Earlier work this paper cites.