Fetching the paper…
Reading the bibliography…
The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training.
The random energy model
B Derrida · 1980
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Automatic learning rate maximization by on-line estimation of the hessian eigenvectors
Yann LeCun, Patrice Y. Simard, and Barak Pearlmutter · 1993
Earlier work this paper cites.
Flat minima
S Hochreiter and J Schmidhuber · 1997
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices, 2005
Jinho Baik, Gérard Ben Arous, and Sandrine Péché · 2005
Earlier work this paper cites.
The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2011
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A Saxe, J McClelland, and S Ganguli · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
On Large-Batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond, 2016
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Hessian-based analysis of large batch training and robustness to adversaries, 2018
Measuring the intrinsic dimension of objective landscapes, 2018
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Later among the works it cites.
The goldilocks zone: Towards better understanding of neural network loss landscapes, 2018
Stanislav Fort and Adam Scherlis · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks, 2018
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density, 2019
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Large scale structure of neural network loss landscapes, 2019
Stanislav Fort and Stanislaw Jastrzebski · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W. Mahoney · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace, 2018
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
On the relation between the sharpest directions of dnn loss and the sgd step length
Stanislaw Jastrzebski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks, 2017a
Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou
Cited in the paper.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians, 2019a
Vardan Papyan
Cited in the paper.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Léon Bottou, and Yann LeCun
Cited in the paper.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan
Cited in the paper.
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Stiffness: A new perspective on generalization in neural networks, 2019
Stanislav Fort, Paweł Krzysztof Nowak, and Srini Narayanan · 2019
Closest in time.
Deep ensembles: A loss landscape perspective
Anonymous · 2020
Closest in time.