Fetching the paper…
Reading the bibliography…
Understanding the loss surface of a neural network is fundamentally important to the understanding of deep learning.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Data Mining: Practical machine learning tools and techniques
Ian H Witten, Eibe Frank, Mark A Hall, and Christopher J Pal · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2017
Earlier work this paper cites.
Global optimality in neural network training
Benjamin D. Haeffele and Rene Vidal · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
A survey on deep learning in medical image analysis
Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen Awm Van Der Laak, Bram Van Ginneken, and Clara I Sánchez · 2017
Earlier work this paper cites.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
No spurious local minima in a two hidden unit relu network
Chenwei Wu, Jiajun Luo, and Jason D Lee · 2018
Later among the works it cites.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Empirical risk landscape analysis for understanding deep neural networks
Pan Zhou and Jiashi Feng · 2018
Later among the works it cites.
Critical points of neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
The multilinear structure of relu networks
Thomas Laurent and James von Brecht · 2018
Cited alongside, same era.
Over-parameterized deep neural networks have no strict local minima for any continuous activations
Dawei Li, Tian Ding, and Ruoyu Sun · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Understanding the loss surface of neural networks for binary classification
Shiyu Liang, Ruoyu Sun, Yixuan Li, and Rayadurgam Srikant · 2018
Cited alongside, same era.
The landscape of empirical risk for nonconvex losses
Song Mei, Yu Bai, Andrea Montanari, et al · 2018
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Quynh Nguyen · 2019
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2019
Later among the works it cites.
SGD converges to global minimum in deep learning via star-convex path
Yi Zhou, Junjie Yang, Huishuai Zhang, Yingbin Liang, and Vahid Tarokh · 2019
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2019
Later among the works it cites.
Truth or backpropaganda? an empirical investigation of deep learning theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 2020
Closest in time.
Elimination of all bad local minima in deep learning
Kenji Kawaguchi and Leslie Pack Kaelbling · 2020
Closest in time.