Fetching the paper…
Reading the bibliography…
We consider the optimization problem associated with training simple ReLU neural networks of the form $\mathbf{x}\mapsto \sum_{i=1}^{k}\max\{0,\mathbf{w}_i^\top \mathbf{x}\}$ with respect to the squared loss.
Local minima and back propagation
T. Poston, C.-N. Lee, Y. Choie, and Y. Kwon · 1991
Earlier work this paper cites.
Exponentially many local minima for single neurons
P. Auer, M. Herbster, and M. K. Warmuth · 1996
Earlier work this paper cites.
The concentration of measure phenomenon
M. Ledoux · 2005
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho and L. K. Saul · 2009
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
B. D. Haeffele and R. Vidal · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
J. Sun, Q. Qu, and J. Wright · 2015
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Cited alongside, same era.
Distribution-specific hardness of learning neural networks
O. Shamir · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
S. S. Du, J. D. Lee, Y. Tian, B. Poczos, and A. Singh · 2017
Closest in time.
Porcupine neural networks:(almost) all local optima are global
S. Feizi, H. Javadi, J. Zhang, and D. Tse · 2017
Closest in time.
Learning one-hidden-layer neural networks with landscape design
R. Ge, J. D. Lee, and T. Ma · 2017
Closest in time.
Convergence analysis of two-layer neural networks with relu activation
Y. Li and Y. Yuan · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Swirszcz, W. M. Czarnecki, and R. Pascanu · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Theoretical properties of the global optimizer of two layer neural network
D. Boob and G. Lan · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Cited alongside, same era.
Q. Nguyen and M. Hein · 2017
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2017
Closest in time.
Y. Tian · 2017
Closest in time.
Electron-proton dynamics in deep learning
Q. Zhang, R. Panigrahy, S. Sachdeva, and A. Rahimi · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon · 2017
Closest in time.