Fetching the paper…
Reading the bibliography…
We introduce a simple algorithm, True Asymptotic Natural Gradient Optimization (TANGO), that converges to a true natural gradient descent in the limit of small learning rates, without explicit Fisher matrix estimation.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Numerical solution of stochastic differential equations
Peter E. Kloeden and Eckhard Platen · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Efficient backprop
Yann Le Cun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Almost sure convergence of two time-scale stochastic approximation algorithms
Vladislav B Tadic · 2004
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Nicolas Le Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2007
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Éric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Learning recurrent neural networks with Hessian-free optimization
James Martens and Ilya Sutskever · 2011
Cited alongside, same era.
Training deep and recurrent neural networks with Hessian-free optimization
James Martens and Ilya Sutskever · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Francis Bach and Eric Moulines · 2013
Cited alongside, same era.
Metric-free natural gradient for joint-training of boltzmann machines
Guillaume Desjardins, Razvan Pascanu, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Yann Ollivier · 2015
Later among the works it cites.
Second order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2016
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2016
Later among the works it cites.
Practical riemannian neural networks
Gaétan Marceau-Caron and Yann Ollivier · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Martens · 2014
Cited alongside, same era.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2015
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Cited alongside, same era.
Two time-scale stochastic approximation with controlled markov noise and off-policy temporal-difference learning
Prasenjit Karmakar and Shalabh Bhatnagar · 2017
Closest in time.
Online natural gradient as a kalman filter
Yann Ollivier · 2017
Closest in time.