Fetching the paper…
Reading the bibliography…
Hessian based measures of flatness, such as the trace, Frobenius and spectral norms, have been argued, used and shown to relate to generalisation.
A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines
Michael F Hutchinson · 1990
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Bayesian methods for adaptive models
David JC MacKay · 1992
Earlier work this paper cites.
Matrices, moments and quadrature
Gene H Golub and Gérard Meurant · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Iterate averaging helps: An alternative perspective in deep learning
Diego Granziol, Xingchen Wan, and Stephen Roberts · 2003
Earlier work this paper cites.
The Lanczos and conjugate gradient algorithms in finite precision arithmetic
Gérard Meurant and Zdeněk Strakoš · 2006
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Matrix computations , volume 3
Gene H Golub and Charles F Van Loan · 2012
Earlier work this paper cites.
Distributions of angles in random packing on spheres
Tony Cai, Jianqing Fan, and Tiefeng Jiang · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Improved bounds on sample size for implicit matrix trace estimators
Farbod Roosta-Khorasani and Uri Ascher · 2015
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Eigenvalues of the Hessian in deep learning: Singularity and beyond
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzkbski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Later among the works it cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving generalization performance by switching from Adam to SGD
Nitish Shirish Keskar and Richard Socher · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Empirical analysis of the Hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Cited alongside, same era.
Deep Frank-Wolfe for neural network optimization
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2018
Cited alongside, same era.
Chiyuan Zhang, Qianli Liao, Alexander Rakhlin, Brando Miranda, Noah Golowich, and Tomaso Poggio · 2018
Later among the works it cites.
An investigation into neural net optimization via Hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
Diego Granziol, Xingchen Wan, Timur Garipov, Dmitry Vetrov, and Stephen Roberts · 2019
Later among the works it cites.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
A simple baseline for Bayesian uncertainty in deep learning
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
A scale invariant flatness measure for deep network minima
Akshay Rangamani, Nam H Nguyen, Abhishek Kumar, Dzung Phan, Sang H Chin, and Trac D Tran · 2019
Later among the works it cites.
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2019
Later among the works it cites.
The break-even point on the optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Closest in time.