Fetching the paper…
Reading the bibliography…
It is well-known that the Hessian of deep loss landscape matters to optimization, generalization, and even robustness of deep learning.
Emergent properties of the local geometry of neural loss landscapes
Fort, S. and Ganguli, S. (2019) · 1910
Earlier work this paper cites.
Non-gaussianity of stochastic gradient noise
Panigrahi, A., Somani, R., Goyal, N., and Netrapalli, P. (2019) · 1910
Earlier work this paper cites.
The kolmogorov-smirnov test for goodness of fit
Massey Jr, F. J. (1951) · 1951
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation. iii
Davis, C. and Kahan, W. M. (1970) · 1970
Earlier work this paper cites.
The principle of maximum entropy
Guiasu, S. and Shenitzer, A. (1985) · 1985
Earlier work this paper cites.
Keeping neural networks simple by minimising the description length of weights
Hinton, G. E. and Camp, D. V. (1993) · 1993
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J. (1995) · 1995
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
The minimum description length principle in coding and modeling
Barron, A., Rissanen, J., and Yu, B. (1998) · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y. (1998) · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
The protein data bank
Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., Shindyalov, I. N., and Bourne, P. E. (2000) · 2000
Earlier work this paper cites.
Anisotropy of fluctuation dynamics of proteins with an elastic network model
Atilgan, A. R., Durell, S., Jernigan, R. L., Demirel, M. C., Keskin, O., and Bahar, I. (2001) · 2001
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Beal, M. J. (2003) · 2003
Earlier work this paper cites.
Tutorial on maximum likelihood estimation
Myung, I. J. (2003) · 2003
Earlier work this paper cites.
Energy-based models for sparse overcomplete representations
Teh, Y. W., Welling, M., Osindero, S., and Hinton, G. E. (2003) · 2003
Earlier work this paper cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J. (2019) · 2003
Earlier work this paper cites.
Problems with fitting to the power-law distribution
Goldstein, M. L., Morris, S. A., and Yen, G. G. (2004) · 2004
Earlier work this paper cites.
Protein tolerance to random amino acid change
Guo, H. H., Choe, J., and Loeb, L. A. (2004) · 2004
Earlier work this paper cites.
Variational learning and bits-back coding: an information-theoretic view to bayesian learning
Honkela, A. and Valpola, H. (2004) · 2004
Earlier work this paper cites.
The lanczos and conjugate gradient algorithms in finite precision arithmetic
Meurant, G. and Strakoš, Z. (2006) · 2006
Earlier work this paper cites.
The minimum description length principle
Grünwald, P. D. (2007) · 2007
Earlier work this paper cites.
Information and complexity in statistical modeling
Rissanen, J. (2007) · 2007
Earlier work this paper cites.
Proteins: coexistence of stability and flexibility
Reuveni, S., Granek, R., and Klafter, J. (2008) · 2008
Earlier work this paper cites.
Power-law distributions in empirical data
Clauset, A., Shalizi, C. R., and Newman, M. E. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Global dynamics of proteins: bridging between structure and function
Bahar, I., Lezon, T. R., Yang, L.-W., and Eyal, E. (2010) · 2010
Earlier work this paper cites.
On the use of stochastic hessian information in optimization methods for machine learning
Byrd, R. H., Chin, G. M., Neveitt, W., and Nocedal, J. (2011) · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Earlier work this paper cites.
A survey of label-noise representation learning: Past, present and future
Han, B., Yao, Q., Liu, T., Niu, G., Tsang, I. W., Kwok, J. T., and Sugiyama, M. (2020) · 2011
Earlier work this paper cites.
Stable weight decay regularization
Xie, Z., Sato, I., and Sugiyama, M. (2020) · 2011
Cited alongside, same era.
Zipf’s law, power laws and maximum entropy
Visser, M. (2013) · 2013
Cited alongside, same era.
powerlaw: a python package for analysis of heavy-tailed distributions
Alstott, J., Bullmore, E., and Plenz, D. (2014) · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
A useful variant of the davis–kahan theorem for statisticians
Yu, Y., Wang, T., and Samworth, R. J. (2015) · 2015
diffgrad: An optimization method for convolutional neural networks
Dubey, S. R., Chakraborty, S., Roy, S. K., Mukherjee, S., Singh, S. K., and Chaudhuri, B. B. (2019) · 2019
Later among the works it cites.
The goldilocks zone: Towards better understanding of neural network loss landscapes
Fort, S. and Scherlis, A. (2019) · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y. (2019) · 2019
Later among the works it cites.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
He, F., Liu, T., and Tao, D. (2019) · 2019
Later among the works it cites.
The asymptotic spectrum of the hessian of dnn throughout training
Jacot, A., Gabriel, F., and Hongler, C. (2019) · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y. (2016) · 2016
Cited alongside, same era.
Learning thermodynamics with boltzmann machines
Torlai, G. and Melko, R. G. (2016) · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Jastrzkebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. (2017) · 2017
Cited alongside, same era.
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S. (2019) · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J. (2019) · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X. (2019) · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Papyan, V. (2019) · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M. (2019) · 2019
Later among the works it cites.
High-dimensional geometry of population responses in visual cortex
Stringer, C., Pachitariu, M., Steinmetz, N., Carandini, M., and Harris, K. D. (2019) · 2019
Later among the works it cites.
Bridging mode connectivity in loss landscapes and adversarial robustness
Zhao, P., Chen, P.-Y., Das, P., Ramamurthy, K. N., and Lin, X. (2019) · 2019
Later among the works it cites.
Statistical mechanics of deep learning
Bahri, Y., Kadmon, J., Pennington, J., Schoenholz, S. S., Sohl-Dickstein, J., and Ganguli, S. (2020) · 2020
Later among the works it cites.
Hessian based analysis of sgd for deep nets: Dynamics and generalization
Li, X., Gu, Q., Zhou, Y., Chen, T., and Banerjee, A. (2020) · 2020
Later among the works it cites.
Dimensional reduction in evolving spin-glass model: correlation of phenotypic responses to environmental and mutational changes
Sakata, A. and Kaneko, K. (2020) · 2020
Later among the works it cites.
Evolutionary dimension reduction in phenotypic space
Sato, T. U. and Kaneko, K. (2020) · 2020
Later among the works it cites.
Functional sensitivity and mutational robustness of proteins
Tang, Q.-Y., Hatakeyama, T. S., and Kaneko, K. (2020) · 2020
Later among the works it cites.
Long-range correlation in protein dynamics: Confirmation by structural data and normal mode analysis
Tang, Q.-Y. and Kaneko, K. (2020) · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Thomas, V., Pedregosa, F., Merriënboer, B., Manzagol, P.-A., Bengio, Y., and Le Roux, N. (2020) · 2020
Later among the works it cites.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Tsuzuku, Y., Sato, I., and Sugiyama, M. (2020) · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S. (2020) · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W. (2020) · 2020
Later among the works it cites.
The heavy-tail phenomenon in sgd
Gurbuzbalaban, M., Simsekli, U., and Zhu, L. (2021) · 2021
Later among the works it cites.
Multiplicative noise and heavy tails in stochastic optimization
Hodgkinson, L. and Mahoney, M. (2021) · 2021
Later among the works it cites.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Li, Z., Malladi, S., and Arora, S. (2021) · 2021
Later among the works it cites.
Hessian eigenspectra of more realistic nonlinear models
Liao, Z. and Mahoney, M. W. (2021) · 2021
Later among the works it cites.
A deeper look at the hessian eigenspectrum of deep neural networks and its applications to regularization
Sankar, A. R., Khasbage, Y., Vigneswaran, R., and Balasubramanian, V. N. (2021) · 2021
Later among the works it cites.
Dynamics-evolution correspondence in protein structures
Tang, Q.-Y. and Kaneko, K. (2021) · 2021
Later among the works it cites.
Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum
Xie, Z., Wang, X., Zhang, H., Sato, I., and Sugiyama, M. (2022) · 2022
Closest in time.