Fetching the paper…
Reading the bibliography…
We study whether a depth two neural network can learn another depth two network using gradient descent.
Mathematical aspects of classical and celestial mechanics
Vladimir I Arnold, Valery V Kozlov, and Anatoly I Neishtadt · 1985
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L. Rivest · 1988
Earlier work this paper cites.
Back propagation fails to separate where perceptrons succeed
Martin L Brady, Raghu Raghavan, and Joseph Slawny · 1989
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K Warmuth · 1996
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivans and Alexander A Sherstov · 2006
Earlier work this paper cites.
Learning kernel-based halfspaces with the 0-1 loss
Shai Shalev-Shwartz, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Later among the works it cites.
Deep residual networks with exponential linear unit
Anish Shah, Eashan Kadam, Hena Shah, Sameer Shinde, and Sandip Shingade · 2016
Later among the works it cites.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Closest in time.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuchen Zhang, Jason D. Lee, Martin J. Wainwright, and Michael I. Jordan · 2015
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Reliably learning the relu in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Closest in time.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Closest in time.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Closest in time.