Fetching the paper…
Reading the bibliography…
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian).
Fast exact multiplication by the Hessian
Pearlmutter, B. A. (1994) · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Matrix Differential Calculus with Applications in Statistics and Econometrics
Magnus, J. R. and Neudecker, H. (1999) · 1999
Earlier work this paper cites.
The curious history of Faà di Bruno’s formula
Johnson, W. P. (2002) · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N. (2002) · 2002
Earlier work this paper cites.
Second-order stagewise backpropagation for Hessian-matrix analyses and investigation of negative curvature
Mizutani, E. and Dreyfus, S. E. (2008) · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J. (2014) · 2014
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R. (2015) · 2015
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers
Grosse, R. and Martens, J. (2016) · 2016
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Feedforward and recurrent neural networks backward propagation and Hessian in matrix form
Naumov, M. (2017) · 2017
Later among the works it cites.
Block-diagonal Hessian-free optimization for training neural networks
Zhang, H., Xiong, C., Bradbury, J., and Socher, R. (2017) · 2017
Later among the works it cites.
The outer product structure of neural network derivatives
Bakker, C., Henry, M. J., and Hodas, N. O. (2018) · 2018
Later among the works it cites.
Automatic differentiation in machine learning: A survey
Baydin, A. G., Pearlmutter, B. A., Radul, A. A., and Siskind, J. M. (2018) · 2018
Later among the works it cites.
BDA-PCH: Block-diagonal approximation of positive-curvature Hessian for training neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Botev, A., Ritter, H., and Barber, D. (2017) · 2017
Cited alongside, same era.
Chen, S.-W., Chou, C.-N., and Chang, E. (2018) · 2018
Later among the works it cites.
DeepOBS: A deep learning optimizer benchmark suite
Schneider, F., Balles, L., and Hennig, P. (2019) · 2019
Closest in time.