Fetching the paper…
Reading the bibliography…
Gradient preconditioning is a key technique to integrate the second-order information into gradients for improving and extending gradient-based learning algorithms.
Limitations of the Empirical Fisher Approximation for Natural Gradient Descent, June 2020
Kunstner, F., Balles, L., and Hennig, P · 1905
Earlier work this paper cites.
Efficient Subsampled Gauss-Newton and Natural Gradient Methods for Training Neural Networks
Ren, Y. and Goldfarb, D · 1906
Earlier work this paper cites.
Stokes, J., Izaac, J., Killoran, N., and Carleo, G · 1909
Earlier work this paper cites.
NGBoost: Natural Gradient Boosting for Probabilistic Prediction
Duan, T., Avati, A., Ding, D. Y., Thai, K. K., Basu, S., Ng, A. Y., and Schuler, A · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing, July 2020
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 1910
Earlier work this paper cites.
PyHessian: Neural Networks Through the Lens of the Hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M · 1912
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal Brain Surgeon
Hassibi, B. and Stork, D. G · 1993
Earlier work this paper cites.
Flat Minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Amari, S.-i · 1998
Earlier work this paper cites.
Scalable Second Order Optimization for Deep Learning
Anil, R., Gupta, V., Koren, T., Regan, K., and Singer, Y · 2002
Earlier work this paper cites.
Fast Curvature Matrix-Vector Products for Second-Order Gradient Descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Practical Quasi-Newton Methods for Training Deep Neural Networks
Goldfarb, D., Ren, Y., and Bahamou, A · 2006
Earlier work this paper cites.
Sketchy Empirical Natural Gradient Methods for Deep Learning
Yang, M., Xu, D., Wen, Z., Chen, M., and Xu, P · 2006
Earlier work this paper cites.
ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning
Yao, Z., Gholami, A., Shen, S., Keutzer, K., and Mahoney, M. W · 2006
Earlier work this paper cites.
Optimization of Graph Neural Networks with Natural Gradient Descent, August 2020
Izadi, M. R., Fang, Y., Stevenson, R., and Lin, L · 2008
Earlier work this paper cites.
Topmoumoute Online Natural Gradient Algorithm
Roux, N. L., Manzagol, P.-a., and Bengio, Y · 2008
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Krylov Subspace Descent for Deep Learning
Vinyals, O. and Povey, D · 2011
Earlier work this paper cites.
The Matrix Cookbook, 2012
Petersen, K. B. and Pedersen, M. S · 2012
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Revisiting Natural Gradient for Deep Networks
Pascanu, R. and Bengio, Y · 2014
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization, August 2015
Dauphin, Y. N., de Vries, H., and Bengio, Y · 2015
Cited alongside, same era.
Scaling Up Natural Gradient by Sparsely Factorizing the Inverse Fisher Matrix
Grosse, R. B. and Salakhutdinov, R · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Noisy Natural Gradient as Variational Inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2018
Later among the works it cites.
Efficient Full-Matrix Adaptive Regularization
Agarwal, N., Bullins, B., Chen, X., Hazan, E., Singh, K., Zhang, C., and Zhang, Y · 2019
Later among the works it cites.
DECOUPLED WEIGHT DECAY REGULARIZATION
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Second-Order Stochastic Optimization for Machine Learning in Linear Time
Agarwal, N., Bullins, B., and Hazan, E · 2017
Cited alongside, same era.
Practical Gauss-Newton Optimisation for Deep Learning
Botev, A., Ritter, H., and Barber, D · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Cited alongside, same era.
Understanding Black-box Predictions via Influence Functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
Krishnan, S., Xiao, Y., and Saurous, R. A · 2017
Cited alongside, same era.
Dangel, F., Kunstner, F., and Hennig, P · 2020
Later among the works it cites.
New Insights and Perspectives on the Natural Gradient Method
Martens, J · 2020
Later among the works it cites.
Continual Deep Learning by Functional Regularisation of Memorable Past
Pan, P., Swaroop, S., Immer, A., Eschenhagen, R., Turner, R. E., and Khan, M. E · 2020
Later among the works it cites.
Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FAC
Ueno, Y., Osawa, K., Tsuji, Y., Naruse, A., and Yokota, R · 2020
Later among the works it cites.
Frantar, E., Kurtic, E., and Alistarh, D · 2021
Later among the works it cites.
{NNGeometry: Easy and Fast Fisher Information Matrices and Neural Tangent Kernels in PyTorch}, 2021
George, T · 2021
Later among the works it cites.
Tensor Normal Training for Deep Learning Models
Ren, Y. and Goldfarb, D · 2021
Later among the works it cites.
SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate Curvature
Tang, Z., Jiang, F., Gong, M., Li, H., Wu, Y., Yu, F., Wang, Z., and Wang, M · 2021
Later among the works it cites.
Chapter 3: Metrics, 2022
Grosse, R · 2022
Later among the works it cites.
Mixed precision algorithms in numerical linear algebra
Higham, N. J. and Mary, T · 2022
Later among the works it cites.
Scalable and Practical Natural Gradient for Large-Scale Deep Learning
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Foo, C.-S., and Yokota, R · 2022
Later among the works it cites.
Deep Neural Network Training with Distributed K-FAC
Pauloski, J. G., Huang, L., Xu, W., Chard, K., Foster, I., and Zhang, Z · 2022
Later among the works it cites.
Sketch-Based Empirical Natural Gradient Methods for Deep Learning
Yang, M., Xu, D., Wen, Z., Chen, M., and Xu, P · 2022
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Ollivier, Y · 2049
Closest in time.
A Natural Policy Gradient
Kakade, S. M · 2073
Closest in time.