Fetching the paper…
Reading the bibliography…
We apply state-of-the-art tools in modern high-dimensional numerical linear algebra to approximate efficiently the spectrum of the Hessian of modern deepnets, with tens of millions of parameters, trained on real data.
An iteration method for the solution of the eigenvalue problem of linear differential and integral operators
Cornelius Lanczos · 1950
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marcenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Moments developments and their application to the electronic charge distribution of d bands
Francois Ducastelle and Françoise Cyrot-Lackmann · 1970
Earlier work this paper cites.
Modified moments for harmonic solids
John C Wheeler and Carl Blumstein · 1972
Earlier work this paper cites.
A maximum-entropy approach to the density of states within the recursion method
I Turek · 1988
Earlier work this paper cites.
Multinomial logistic regression algorithm
Dankmar Böhning · 1992
Earlier work this paper cites.
Maximum entropy approach for linear scaling in the electronic structure problem
David A Drabold and Otto F Sankey · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient backprop in neural networks: Tricks of the trade (orr, g. and müller, k., eds.)
Y LeCun, L Bottou, G Orr, and K Muller · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Clusterjob: An automated system for painless and reproducible massive computational experiments
H. Monajemi and D. L. Donoho · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Closest in time.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Closest in time.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten · 2018
Closest in time.
Charles H Martin and Michael W Mahoney · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Approximating spectral densities of large matrices
Lin Lin, Yousef Saad, and Chao Yang · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Making massive computational experiments painless
H. Monajemi, D. L. Donoho, and V. Stodden · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Closest in time.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Closest in time.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Closest in time.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2018
Closest in time.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Closest in time.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Anonymous · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Ambitious data science can be painless
H. Monajemi, R. Murri, E. Yonas, P. Liang, V. Stodden, and D.L. Donoho · 2019
Closest in time.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Closest in time.