New insights and perspectives on the natural gradient method
Original
Martens, J · 2014
Later among the works it cites.
Parallel training of DNNs with natural gradient and parameter averaging
Original
Povey, D., Zhang, X., and Khudanpur, S · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Ollivier, Y · 2015
Later among the works it cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Later among the works it cites.
Distributed second-order optimization using Kronecker-factored approximations
Ba, J., Grosse, R., and Martens, J · 2017
Later among the works it cites.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Later among the works it cites.
Fast approximate natural gradient descent in a Kronecker factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Later among the works it cites.
Limitations of the empirical Fisher approximation for natural gradient descent
Kunstner, F., Hennig, P., and Balles, L · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.