Fetching the paper…
Reading the bibliography…
Numerous researchers recently applied empirical spectral analysis to the study of modern deep learning classifiers.
Freeze and chaos for dnns: an ntk view of batch normalization, checkerboard and boundary effects
Arthur Jacot, Franck Gabriel, and Clément Hongler · 1907
Earlier work this paper cites.
The asymptotic spectrum of the hessian of dnn throughout training
Arthur Jacot, Franck Gabriel, and Clément Hongler · 1910
Earlier work this paper cites.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 1910
Earlier work this paper cites.
An iteration method for the solution of the eigenvalue problem of linear differential and integral operators
Cornelius Lanczos · 1950
Earlier work this paper cites.
Moments developments and their application to the electronic charge distribution of d bands
Francois Ducastelle and Françoise Cyrot-Lackmann · 1970
Earlier work this paper cites.
Modified moments for harmonic solids
John C Wheeler and Carl Blumstein · 1972
Earlier work this paper cites.
A maximum-entropy approach to the density of states within the recursion method
I Turek · 1988
Earlier work this paper cites.
Multinomial logistic regression algorithm
Dankmar Böhning · 1992
Earlier work this paper cites.
Maximum entropy approach for linear scaling in the electronic structure problem
David A Drabold and Otto F Sankey · 1993
Earlier work this paper cites.
Approximation with kronecker products
Charles F Van Loan and Nikos Pitsianis · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Efficient backprop in neural networks: Tricks of the trade (orr, g. and müller, k., eds.)
Y LeCun, L Bottou, G Orr, and K Muller · 1998
Earlier work this paper cites.
A gentle hessian for efficient gradient descent
Ronan Collobert and Samy Bengio · 2004
Earlier work this paper cites.
Applied MANOVA and discriminant analysis , volume 498
Carl J Huberty and Stephen Olejnik · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Scaling up natural gradient by sparsely factorizing the inverse fisher matrix
Roger Grosse and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Clusterjob: An automated system for painless and reproducible massive computational experiments
H. Monajemi and D. L. Donoho · 2015
Earlier work this paper cites.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2016
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Approximating spectral densities of large matrices
Lin Lin, Yousef Saad, and Chao Yang · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Emergent properties of the local geometry of neural loss landscapes
Stanislav Fort and Surya Ganguli · 2019
Later among the works it cites.
Large scale structure of neural network loss landscapes
Stanislav Fort and Stanisław Jastrzębski · 2019
Later among the works it cites.
The goldilocks zone: Towards better understanding of neural network loss landscapes
Stanislav Fort and Adam Scherlis · 2019
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Making massive computational experiments painless
H. Monajemi, D. L. Donoho, and V. Stodden · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Fast estimation of tr(f(a)) via stochastic lanczos quadrature
Shashanka Ubaru, Jie Chen, and Yousef Saad · 2017
Cited alongside, same era.
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Gpu accelerated sub-sampled newton’s method for convex classification problems
Sudhir Kylasa, Fred Roosta, Michael W Mahoney, and Ananth Grama · 2019
Later among the works it cites.
Hessian based analysis of sgd for deep nets: Dynamics and generalization
Xinyan Li, Qilong Gu, Yingxue Zhou, Tiancong Chen, and Arindam Banerjee · 2019
Later among the works it cites.
Traditional and heavy tailed self regularization in neural network models
Michael Mahoney and Charles Martin · 2019
Later among the works it cites.
Ambitious data science can be painless
H. Monajemi, R. Murri, E. Yonas, P. Liang, V. Stodden, and D.L. Donoho · 2019
Later among the works it cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Samet Oymak, Zalan Fabian, Mingchen Li, and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan · 2019
Later among the works it cites.
Evolution of eigenvalue decay in deep networks
Lukas Pfahler and Katharina Morik · 2019
Later among the works it cites.
Information matrices and generalization
Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Mangazol, Yoshua Bengio, and Nicolas Le Roux · 2019
Later among the works it cites.
Eigendamage: Structured pruning in the kronecker-factored eigenbasis
Chaoqi Wang, Roger Grosse, Sanja Fidler, and Guodong Zhang · 2019
Later among the works it cites.
Utilizing second order information in minibatch stochastic variance reduced proximal iterations
Jialei Wang and Tong Zhang · 2019
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
Peng Xu, Fred Roosta, and Michael W Mahoney · 2019
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael Mahoney · 2019
Later among the works it cites.
Chiyuan Zhang, Samy Bengio, and Yoram Singer · 2019
Later among the works it cites.
Asymptotics of wide convolutional neural networks
Anders Andreassen and Ethan Dyer · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Closest in time.
The break-even point on optimization trajectories of deep neural networks
Stanisław Jastrzębski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Inefficiency of k-fac for large batch size training
Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney · 2020
Closest in time.
Anatomy of catastrophic forgetting: Hidden representations and task semantics
Vinay V Ramasesh, Ethan Dyer, and Maithra Raghu · 2020
Closest in time.
Second-order optimization for non-convex machine learning: An empirical study
Peng Xu, Fred Roosta, and Michael W Mahoney · 2020
Closest in time.