Fetching the paper…
Reading the bibliography…
Hessian captures important properties of the deep neural network loss landscape.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Karakida, R., Akaho, S., and Amari, S.-i · 1910
Earlier work this paper cites.
The design of suboptimal linear time-varying systems
Kleinman, D. and Athans, M · 1968
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
McAllester, D. A · 1999
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Heskes, T · 2000
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Laurent, B. and Massart, P · 2000
Earlier work this paper cites.
Bounds for averaging classifiers
Langford, J. and Seeger, M · 2001
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Torralba, A., Fergus, R., and Freeman, W. T · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Optimization for machine learning
Sra, S., Nowozin, S., and Wright, S. J · 2012
Earlier work this paper cites.
A short note on the tail bound of wishart distribution
Zhu, S · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M · 2017
Emergent properties of the local geometry of neural loss landscapes
Fort, S. and Ganguli, S · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Jastrzebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. J · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Papyan, V · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Cited alongside, same era.
pytorch-hessian-eigentings: efficient pytorch hessian eigendecomposition, 2018
Golmant, N., Yao, Z., Gholami, A., Mahoney, M., and Gonzalez, J · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Understanding impacts of high-order loss approximations and features in deep learning interpretation
Singla, S., Wallace, E., Feng, S., and Feizi, S · 2019
Later among the works it cites.
Chain rules for hessian and higher derivatives made easy by tensor calculus
Skorski, M · 2019
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M · 2019
Later among the works it cites.
Modular block-diagonal curvature approximations for feedforward architectures
Dangel, F., Harmeling, S., and Hennig, P · 2020
Closest in time.
The asymptotic spectrum of the hessian of DNN throughout training
Jacot, A., Gabriel, F., and Hongler, C · 2020
Closest in time.
Hessian based analysis of sgd for deep nets: Dynamics and generalization
Li, X., Gu, Q., Zhou, Y., Chen, T., and Banerjee, A · 2020
Closest in time.
Traces of class/cross-class structure pervade deep learning spectra
Papyan, V · 2020
Closest in time.
Hessian eigenspectra of more realistic nonlinear models
Liao, Z. and Mahoney, M. W · 2021
Closest in time.
Analytic insights into structure and rank of neural network hessian maps
Singh, S. P., Bachmann, G., and Hofmann, T · 2021
Closest in time.