Fetching the paper…
Reading the bibliography…
We study the improper learning of multi-layer neural networks.
Training a 3-node neural network is NP-complete
A. L. Blum and R. L. Rivest · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Convex neural networks
Y. Bengio, N. L. Roux, P. Vincent, O. Delalleau, and P. Marcotte · 2005
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
A. R. Klivans, A. Sherstov, et al · 2006
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
S. M. Kakade, K. Sridharan, and A. Tewari · 2009
Earlier work this paper cites.
Learning kernel-based halfspaces with the 0-1 loss
S. Shalev-Shwartz, O. Shamir, and K. Sridharan · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Provable bounds for learning some deep representations
S. Arora, A. Bhaskara, R. Ge, and T. Ma · 2013
Cited alongside, same era.
Random matrices and complexity of spin glasses
A. Auffinger, G. B. Arous, and J. Černỳ · 2013
Cited alongside, same era.
Improving deep neural networks for lvcsr using rectified linear units and dropout
G. E. Dahl, T. N. Sainath, and G. E. Hinton · 2013
Cited alongside, same era.
A fast and accurate dependency parser using neural networks
D. Chen and C. D. Manning · 2014
Later among the works it cites.
The loss surface of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Later among the works it cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Later among the works it cites.
Provable methods for training neural networks with sparse connectivity
H. Sedghi and A. Anandkumar · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Building high-level features using large scale unsupervised learning
Q. V. Le · 2013
Cited alongside, same era.
Sparse activity and sparse connectivity in supervised learning
M. Thom and G. Palm · 2013
Cited alongside, same era.
http://www.iro.umontreal.ca/~lisa/twiki/bin/view.cgi/Public/MnistVariations
Variations on the MNIST digits
Cited in the paper.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Later among the works it cites.
Generalization bounds for neural networks through tensor factorization
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Closest in time.