Fetching the paper…
Reading the bibliography…
Many aspects of the geometry of loss functions in deep learning remain mysterious.
On connected sublevel sets in deep learning
Quynh Nguyen · 1901
Earlier work this paper cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge · 1906
Earlier work this paper cites.
On the convergence of projected gradient processes to singular critical points
J. C. Dunn · 1987
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Manifolds and differential geometry
Jeffrey M. Lee · 2009
Earlier work this paper cites.
Nonlinear programming
Dimitri P. Bertsekas · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Over-parameterized deep neural networks have no strict local minima for any continuous activations
Dawei Li, Tian Ding, and Ruoyu Sun · 2018
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Later among the works it cites.
Spurious valleys in two-layer neural network optimization landscapes
Luca Venturi, Afonso S. Bandeira, and Joan Bruna · 2018
Later among the works it cites.
A critical view of global optimality in deep learning
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Sub-optimal local minima exist for almost all over-parameterized neural networks
Tian Ding, Dawei Li, and Ruoyu Sun · 2019
Later among the works it cites.