Fetching the paper…
Reading the bibliography…
In this paper we study the problem of learning a shallow artificial neural network that best fits a training data set.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Training a 3-node neural network is NPs-complete
A. Blum and R. L. Rivest · 1988
Earlier work this paper cites.
Local minima and back propagation
T. Poston, C-N. Lee, Y. Choie, and Y. Kwon · 1991
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
A. R. Barron · 1994
Earlier work this paper cites.
On-line learning in soft committee machines
D. Saad and S. A. Solla · 1995
Earlier work this paper cites.
Transient dynamics of on-line learning in two-layered neural networks
M. Biehl, P. Riegler, and C. Wohler · 1996
Earlier work this paper cites.
Functional optimization of online algorithms in multilayer neural networks
R. Vicente and N. Caticha · 1997
Earlier work this paper cites.
The concentration of measure phenomenon
Michel Ledoux · 2005
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
An introduction to random matrices
G. W. Anderson, A. Guionnet, and O. Zeitouni · 2009
Earlier work this paper cites.
The isotron algorithm: High-dimensional isotonic regression
A. T. Kalai and R. Sastry · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Restricted isometry property of matrices with independent columns and neighborly polytopes by random sampling
R. Adamczak, A. E. Litvak, A. Pajor, and N. Tomczak-Jaegermann · 2011
Earlier work this paper cites.
Sharp bounds on the rate of convergence of the empirical covariance matrix
Radosław Adamczak, Alexander E Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann · 2011
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
S. M. Kakade, V. Kanade, O. Shamir, and A. Kalai · 2011
Earlier work this paper cites.
Learning kernel-based halfspaces with the 0-1 loss
S. Shalev-Shwartz, O. Shamir, and K. Sridharan · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Acoustic modeling using deep belief networks
A. Mohamed, G. E. Dahl, and G. Hinton · 2012
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Cited alongside, same era.
Algorithms and Theory for Clustering and Nonconvex Quadratic Programming
M. Soltanolkotabi · 2014
Cited alongside, same era.
A note on the hanson-wright inequality for random vectors with dependencies
R. Adamczak · 2015
Cited alongside, same era.
Phase retrieval via Wirtinger flow: Theory and algorithms
E. J. Candes, X. Li, and M. Soltanolkotabi · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Diversity leads to generalization in neural networks
B. Xie, Y. Liang, and L. Song · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Y. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
L1-regularized neural networks are improperly learnable in polynomial time
Y. Zhang, J. D. Lee, and M. I. Jordan · 2016
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Closest in time.
Optimal approximation with sparsely connected deep neural networks
H. Bolcskei, P. Grohs, G. Kutyniok, and P. Petersen · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Global optimality in tensor factorization, deep learning, and beyond
B. D. Haeffele and R. Vidal · 2015
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Cited alongside, same era.
Finding approximate local minima for nonconvex optimization in linear time
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2016
Cited alongside, same era.
Gradient descent efficiently finds the cubic-regularized non-convex newton step
Y. Carmon and J. C. Duchi · 2016
Cited alongside, same era.
Accelerated methods for non-convex optimization
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2016
Cited alongside, same era.
Closest in time.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Closest in time.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Closest in time.
The CNN as a guided multilayer RECOS transform [lecture notes]
C. J. Kuo · 2017
Closest in time.
Algorithmic regularization in over-parameterized matrix recovery
Y. Li, T. Ma, and H. Zhang · 2017
Closest in time.
Convergence analysis of two-layer neural networks with ReLU activation
Y. Li and Y. Yuan · 2017
Closest in time.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
I. Safran and O. Shamir · 2017
Closest in time.
Learning ReLUs via gradient descent
M. Soltanolkotabi · 2017
Closest in time.
M. Soltanolkotabi · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Closest in time.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Y. Tian · 2017
Closest in time.
Energy propagation in deep convolutional neural networks
T. Wiatowski, P. Grohs, and H. Bolcskei · 2017
Closest in time.
Electron-proton dynamics in deep learning
Q. Zhang, R. Panigrahy, S. Sachdeva, and A. Rahimi · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon · 2017
Closest in time.