Fetching the paper…
Reading the bibliography…
Loss landscape analysis is extremely useful for a deeper understanding of the generalization ability of deep neural network models.
Zhang, C.; Bengio, S.; and Singer, Y. 2019 · 1902
Earlier work this paper cites.
Deep Ensembles: A Loss Landscape Perspective
Fort, S.; Hu, H.; and Lakshminarayanan, B. 2019 · 1912
Earlier work this paper cites.
An iteration method for the solution of the eigenvalue problem of linear differential and integral operators
Lanczos, C. 1950 · 1950
Earlier work this paper cites.
Neural networks and principal component analysis: learning from examples without local minima
Baldi, P.; and Hornik, K. 1988 · 1988
Earlier work this paper cites.
Numerical data on neocortical neurons in adult rat, with special reference to the GABA population
Beaulieu, C. 1993 · 1993
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, P.; Warmuth, M.; and K, M. K. 1996 · 1996
Earlier work this paper cites.
Flat Minima
Hochreiter, Schmidhuber, a. J. 1997 · 1997
Earlier work this paper cites.
Optimal transport – Old and new , xxii+973
Villani, C. 2008 · 2008
Earlier work this paper cites.
Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix
Avron, H.; and Toledo, S. 2011 · 2011
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N.; Pascanu, R.; Gulcehre, C.; Cho, K.; Ganguli, S.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Open Problem: The landscape of the loss surfaces of multilayer networks
Anna, C.; LeCun, Y.; and Arous, G. B. 2015 · 2015
Cited alongside, same era.
Optimizing Neural Networks with Kronecker-Factored Approximate Curvature
Martens, J.; and Grosse, R. 2015 · 2015
Cited alongside, same era.
Deep Learning without Poor Local Minima
Kawaguchi, K. 2016 · 2016
Cited alongside, same era.
Approximating spectral densities of large matrices
Lin, L.; Saad, Y.; and Yang, C. 2016 · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L.; Bottou, L.; and LeCun, Y. 2016 · 2016
Cited alongside, same era.
Practical Gauss-Newton Optimisation for Deep Learning
Botev, A.; Ritter, H.; and Barber, D. 2017 · 2017
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T.; Izmailov, P.; Podoprikhin, D.; Vetrov, D. P.; and Wilson, A. G. 2018 · 2018
Later among the works it cites.
Gradient Descent Happens in a Tiny Subspace
Gur-Ari, G.; Roberts, D. A.; and Dyer, E. 2018 · 2018
Later among the works it cites.
Visualizing the Loss Landscape of Neural Nets
Li, H.; Xu, Z.; Taylor, G.; Studer, C.; and Goldstein, T. 2018 · 2018
Later among the works it cites.
The Full Spectrum of Deep Net Hessians At Scale: Dynamics with Sample Size
Papyan, V. 2018 · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik, C.; and Stefano, S. 2018 · 2018
Later among the works it cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Chaudhari, P.; Choromanska, A.; Soatto, S.; LeCun, Y.; Baldassi, C.; Borgs, C.; Chayes, J.; Sagun, L.; and Zecchina, R. 2017 · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
Dinh, L.; Pascanu, R.; Bengio, S.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Identity Matters in Deep Learning
Hardt, M.; and Ma, T. 2017 · 2017
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S.; Mudigere, D.; Nocedal, J.; Smelyanskiy, M.; and Tang, P. T. P. 2017 · 2017
Cited alongside, same era.
Sagun, L.; Evci, U.; Güney, V. U.; Dauphin, Y. N.; and Bottou, L. 2018 · 2018
Later among the works it cites.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
Ghorbani; Shankar, K.; and Xiao, Y. 2019 · 2019
Later among the works it cites.
On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
Jastrzebski, S.; Kenton, Z.; Ballas, N.; Fischer, A.; Bengio, Y.; and Storkey, A. 2019 · 2019
Later among the works it cites.
Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
Papyan, V. 2019 · 2019
Later among the works it cites.