Fetching the paper…
Reading the bibliography…
Curvature in form of the Hessian or its generalized Gauss-Newton (GGN) approximation is valuable for algorithms that rely on a local model for the loss to train, compress, or explain deep networks.
Automatic learning rate maximization by on-line estimation of the hessian's eigenvectors
LeCun, Y., Simard, P., and Pearlmutter, B · 1993
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-i · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
On the use of stochastic Hessian information in optimization methods for machine learning
Byrd, R. H., Chin, G. M., Neveitt, W., and Nocedal, J · 2011
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers, 2016
Grosse, R. and Martens, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Earlier work this paper cites.
Eigenvalues of the Hessian in deep learning: Singularity and beyond, 2017
Sagun, L., Bottou, L., and LeCun, Y · 2017
Earlier work this paper cites.
Block-diagonal Hessian-free optimization for training neural networks, 2017
Zhang, H., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Estimating the spectral density of large implicit matrices, 2018
Adams, R. P., Pennington, J., Johnson, M. J., Smith, J., Ovadia, Y., Patton, B., and Saunderson, J · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace, 2018
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Proper ResNet implementation for CIFAR10/CIFAR100 in PyTorch
Idelbayev, Y · 2018
Cited alongside, same era.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Cited alongside, same era.
Empirical analysis of the Hessian of over-parametrized neural networks, 2018
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2018
JAX: composable transformations of Python + NumPy programs
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2020
Later among the works it cites.
BackPACK: Packing more into backprop
Dangel, F., Kunstner, F., and Hennig, P · 2020
Later among the works it cites.
On the promise of the stochastic generalized Gauss-Newton method for training DNNs, 2020
Gargiani, M., Zanelli, A., Diehl, M., and Hutter, F · 2020
Later among the works it cites.
Being Bayesian, even just a bit, fixes overconfidence in ReLU networks
Kristiadi, A., Hein, M., and Hennig, P · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method, 2020
Martens, J · 2020
Later among the works it cites.
WoodFisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An investigation into neural net optimization via Hessian eigenvalue density, 2019
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Cited alongside, same era.
Limitations of the empirical Fisher approximation for natural gradient descent
Kunstner, F., Hennig, P., and Balles, L · 2019
Cited alongside, same era.
Large-scale distributed second-order optimization using Kronecker-factored approximate curvature for deep convolutional neural networks
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
DeepOBS: A deep learning optimizer benchmark suite
Schneider, F., Balles, L., and Hennig, P · 2019
Cited alongside, same era.
PyHessian: Neural networks through the lens of the Hessian, 2019
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M · 2019
Cited alongside, same era.
Later among the works it cites.
Fast approximation of the Gauss–Newton Hessian matrix for the multilayer perceptron
Chen, C., Reiz, S., Yu, C. D., Bungartz, H.-J., and Biros, G · 2021
Closest in time.
Laplace redux–effortless Bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P · 2021
Closest in time.
Deep curvature suite, 2021
Granziol, D., Wan, X., and Garipov, T · 2021
Closest in time.
Descending through a crowded valley – benchmarking deep learning optimizers, 2021
Schmidt, R. M., Schneider, F., and Hennig, P · 2021
Closest in time.
Cockpit: A practical debugging tool for training deep neural networks, 2021
Schneider, F., Dangel, F., and Hennig, P · 2021
Closest in time.