Fetching the paper…
Reading the bibliography…
Neural networks provide a rich class of high-dimensional, non-convex optimization problems.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Symmetric tensors and symmetric tensor rank
Pierre Comon, Gene Golub, Lek-Heng Lim, and Bernard Mourrain · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
On the low-rank approach for semidefinite programs arising in synchronization and community detection
Afonso S Bandeira, Nicolas Boumal, and Vladislav Voroninski · 2016
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Earlier work this paper cites.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Local minima in training of neural networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Cited alongside, same era.
Geometrical insights for implicit generative modeling
Leon Bottou, Martin Arjovsky, David Lopez-Paz, and Maxime Oquab · 2017
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Closest in time.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A Hamprecht · 2018
Closest in time.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Closest in time.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Closest in time.
Characterizing implicit bias in terms of optimization geometry
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simon S Du, Jason D Lee, Yuandong Tian, Barnabas Poczos, and Aarti Singh · 2017
Cited alongside, same era.
Porcupine neural networks:(almost) all local optima are global
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Cited alongside, same era.
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Closest in time.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Closest in time.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Closest in time.
How does batch normalization help optimization?(no, it is not about internal covariate shift)
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Closest in time.
Are resnets provably better than linear predictors?
Ohad Shamir · 2018
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Closest in time.
Chao Ma, Lei Wu, et al · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Closest in time.