Fetching the paper…
Reading the bibliography…
In optimization for Machine learning (ML), it is typical that curvature-matrix (CM) estimates rely on an exponential average (EA) of local estimates (giving EA-CM algorithms).
LeCun, Y.; Bottou, L.; Orr, G.; Muller, K. Efficient backprop. Neural networks: Tricks of the trade, pages 546-546 (1998)
1998
Earlier work this paper cites.
Amari, S. I. Natural gradient works efficiently in learning, Neural Computation, 10(20), pp. 251-276 (1998)
1998
Earlier work this paper cites.
Park, H.; Amari, S.-I.; Fukumizu, K. Adaptive natural gradient learning algorithms for various stochastic models. Neural Networks, 13(7):755-764 (2000)
2000
Earlier work this paper cites.
Schaul, T.; Zhang, S.; LeCun, Y. No more pesky learning rates. In ICML (2013)
2013
Earlier work this paper cites.
2015
Cited alongside, same era.
Ba, J.; Kingma, D. Adam: A method for stochastic optimization, ICLR (2015)
2015
Cited alongside, same era.
Wu, Y.; Mansimov, E.; Grosse, R. B.; Liao, S.; Ba, J. Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation. In Advances in neural information processing systems, pages 5285-5294 (2017)
2017
Cited alongside, same era.
Bottou, L.; Curtis, F. E.; Nocedal, J. Optimization methods for large-scale machine learning (2018)
2018
Cited alongside, same era.
Martens, J. New insights and perspectives on the natural gradient method, arXiv:1412.1193 (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…