Fetching the paper…
Reading the bibliography…
We analyze the convergence of (stochastic) gradient descent algorithm for learning a convolutional filter with Rectified Linear Unit (ReLU) activation function.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
Jiří Šíma · 2002
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Earlier work this paper cites.
An adaptive accelerated proximal gradient method and its homotopy continuation for sparse optimization
Qihang Lin and Lin Xiao · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Earlier work this paper cites.
Learning halfspaces and neural networks with random initialization
Yuchen Zhang, Jason D Lee, Martin J Wainwright, and Michael I Jordan · 2015
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Cited alongside, same era.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
Antoine Gautier, Quynh N Nguyen, and Matthias Hein · 2016
Cited alongside, same era.
Reliably learning the relu in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Closest in time.
Gradient descent can take exponential time to escape saddle points
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Barnabas Poczos, and Aarti Singh · 2017
Closest in time.
Learning depth-three neural networks in polynomial time
Surbhi Goel and Adam Klivans · 2017
Closest in time.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Cited alongside, same era.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi · 2016
Cited alongside, same era.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Cited alongside, same era.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2016
Cited alongside, same era.
Closest in time.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Closest in time.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Closest in time.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Closest in time.
Yuandong Tian · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Closest in time.
The landscape of deep learning algorithms
Pan Zhou and Jiashi Feng · 2017
Closest in time.