Fetching the paper…
Reading the bibliography…
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Tom Heskes · 2000
Earlier work this paper cites.
Adaptive natural gradient learning algorithms for various stochastic models
Hyeyoung Park, Shun-ichi Amari, and Kenji Fukumizu · 2000
Earlier work this paper cites.
SciPy: Open source scientific tools for Python, 2001
Eric Jones, Travis Oliphant, Pearu Peterson, et al · 2001
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N. Schraudolph · 2002
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Nicolas Le Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2007
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Jan Peters, and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Efficient natural evolution strategies
Yi Sun, Daan Wierstra, Tom Schaul, and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
A fast natural Newton method
Nicolas Le Roux and Andrew Fitzgibbon · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Fisher scoring: An interpolation family and its Monte Carlo implementations
Yong Wang · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a Kronecker-factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Later among the works it cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zac Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Amos Storkey, and Yoshua Bengio · 2018
Later among the works it cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Mohammad Emtiyaz Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, and Akash Srivastava · 2018
Later among the works it cites.
SLANG: Fast structured covariance approximations for Bayesian deep learning with natural gradient
Aaron Mishkin, Frederik Kunstner, Didrik Nielsen, Mark Schmidt, and Mohammad Emtiyaz Khan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Riemannian metrics for neural networks I: feedforward networks
Yann Ollivier · 2015
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Arnold Salas, Stefan Zohren, and Stephen Roberts · 2018
Later among the works it cites.
Noisy natural gradient as variational inference
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2018
Later among the works it cites.
Universal statistics of fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Closest in time.
Large-scale distributed second-order optimization using Kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Closest in time.
Information matrices and generalization
Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2019
Closest in time.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Closest in time.
Approximate Fisher information matrix to characterise the training of deep neural networks
Zhibin Liao, Tom Drummond, Ian Reid, and Gustavo Carneiro · 2020
Closest in time.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2020
Closest in time.