Training a 3-node neural network is NP-complete
Blum, Avrim and Rivest, Ronald L · 1989
Earlier work this paper cites.
Efficient backprop
LeCun, Yann, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 1998
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
Šíma, Jiří · 2002
Earlier work this paper cites.
Kernel methods for deep learning
Cho, Youngmin and Saul, Lawrence K · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
The noisy power method: A meta algorithm with applications
Hardt, Moritz and Price, Eric · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, Roi, Shalev-Shwartz, Shai, and Shamir, Ohad · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
Original
Sedghi, Hanie and Anandkumar, Anima · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2015
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Ge, Rong, Huang, Furong, Jin, Chi, and Yuan, Yang · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Original
Haeffele, Benjamin D and Vidal, René · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Original
Janzamin, Majid, Sedghi, Hanie, and Anandkumar, Anima · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Neyshabur, Behnam, Salakhutdinov, Ruslan R, and Srebro, Nati · 2015
Earlier work this paper cites.