Fetching the paper…
Reading the bibliography…
The Hessian of neural networks can be decomposed into a sum of two matrices: (i) the positive semidefinite generalized Gauss-Newton matrix G, and (ii) the matrix H containing negative eigenvalues.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X. J · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R. B · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Singularity of the hessian in deep learning
Sagun, L., Bottou, L., and LeCun, Y · 2016
Cited alongside, same era.
Escaping saddles with stochastic gradients
Daneshmand, H., Kohler, J., Lucchi, A., and Hofmann, T · 2018
Later among the works it cites.
Shampoo: Preconditioned stochastic tensor optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Later among the works it cites.
The benefits of over-parameterization at initialization in deep relu networks
Arpit, D. and Bengio, Y · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…