Fetching the paper…
Reading the bibliography…
We provide new theoretical insights on why over-parametrization is effective in learning neural networks.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Local minima and back propagation
Timothy Poston, C-N Lee, Y Choie, and Yonghoon Kwon · 1991
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
On the rank of extreme matrices in semidefinite programs and the multiplicity of optimal eigenvalues
Gábor Pataki · 1998
Earlier work this paper cites.
The geometry of semidefinite programming
Gábor Pataki · 2000
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi S Jaakkola · 2005
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Structured low-rank matrix factorization: Optimality, algorithm, and applications to image processing
Benjamin Haeffele, Eric Young, and Rene Vidal · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Cited alongside, same era.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Later among the works it cites.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Nick Harvey, Chris Liaw, and Abbas Mehrabian · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix recovery
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Later among the works it cites.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2017
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2017
Later among the works it cites.
Towards provable learning of polynomial neural networks using low-rank matrix estimation
Mohammadreza Soltani and Chinmay Hegde · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Later among the works it cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Daniel Soudry and Elad Hoffer · 2017
Later among the works it cites.
Yuandong Tian · 2017
Later among the works it cites.
Energy propagation in deep convolutional neural networks
Thomas Wiatowski, Philipp Grohs, and Helmut Bölcskei · 2017
Later among the works it cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Later among the works it cites.
The landscape of deep learning algorithms
Pan Zhou and Jiashi Feng · 2017
Later among the works it cites.
Smoothed analysis for low-rank solutions to semidefinite programs in quadratic penalty form
Srinadh Bhojanapalli, Nicolas Boumal, Prateek Jain, and Praneeth Netrapalli · 2018
Closest in time.
Generalization error bounds for noisy, iterative algorithms
Ankit Pensia, Varun Jog, and Po-Ling Loh · 2018
Closest in time.