Fetching the paper…
Reading the bibliography…
We cast Amari's natural gradient in statistical learning as a specific case of Kalman filtering.
Information theory and statistics
Solomon Kullback · 1968
Earlier work this paper cites.
Stochastic processes and filtering theory
Andrew H. Jazwinski · 1970
Earlier work this paper cites.
Theory and Practice of Recursive Identification
Lennart Ljung and Torsten Söderström · 1983
Earlier work this paper cites.
Riemannian geometry
S. Gallot, D. Hulin, and J. Lafontaine · 1987
Earlier work this paper cites.
Training multilayer perceptrons with the extended Kalman algorithm
Sharad Singhal and Lance Wu · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Comparative analysis of backpropagation and the extended Kalman filter for training multilayer perceptrons
Dennis W. Ruck, Steven K. Rogers, Matthew Kabrisky, Peter S. Maybeck, and Mark E. Oxley · 1992
Earlier work this paper cites.
Training recurrent networks using the extended Kalman filter
Ronald J Williams · 1992
Earlier work this paper cites.
Methods of information geometry
Shun-ichi Amari and Hiroshi Nagaoka · 1993
Earlier work this paper cites.
Incremental least squares methods and the extended Kalman filter
Dimitri P. Bertsekas · 1996
Earlier work this paper cites.
Dual Kalman filtering methods for nonlinear prediction, smoothing and estimation
Eric A. Wan and Alex T. Nelson · 1996
Earlier work this paper cites.
Convergence analysis of the extended Kalman filter used as an observer for nonlinear deterministic discrete-time systems
M. Boutayeb, H. Rafaralahy, and M. Darouach · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Cited alongside, same era.
Hierarchical Bayesian models for regularization in sequential learning
João FG de Freitas, Mahesan Niranjan, and Andrew H. Gee · 2000
Cited alongside, same era.
Asymptotic statistics
A.W. van der Vaart · 2000
Cited alongside, same era.
Kalman filtering and neural networks
Simon Haykin · 2001
Cited alongside, same era.
Filtering, predictive, and smoothing Cramér–Rao bounds for discrete-time nonlinear dynamic systems
Miroslav Šimandl, Jakub Královec, and Petr Tichavskỳ · 2001
Cited alongside, same era.
Tutorial on training recurrent neural networks, covering BPTT, RTRL, EKF and the “echo state network” approach
Herbert Jaeger · 2002
Bayesian filtering and smoothing
Simo Särkkä · 2013
Later among the works it cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Later among the works it cites.
Kalman filtering: Theory and practice using MATLAB
Mohinder S. Grewal and Angus P. Andrews · 2015
Later among the works it cites.
Scaling up natural gradient by sparsely factorizing the inverse Fisher matrix
Roger B. Grosse and Ruslan Salakhutdinov · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger B. Grosse · 2015
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Yann Ollivier · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large scale online learning
Léon Bottou and Yann LeCun · 2003
Cited alongside, same era.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Pattern recognition and machine learning
Christopher M. Bishop · 2006
Cited alongside, same era.
Optimal state estimation: Kalman, H ∞ H_{\infty} , and nonlinear approaches
Dan Simon · 2006
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Nicolas Le Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2007
Cited alongside, same era.
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Later among the works it cites.
Training recurrent networks online without backtracking
Yann Ollivier, Corentin Tallec, and Guillaume Charpiat · 2015
Later among the works it cites.
Practical Riemannian neural networks
Gaétan Marceau-Caron and Yann Ollivier · 2016
Later among the works it cites.
Kalman-based stochastic gradient method with stop condition and insensitivity to conditioning
Vivak Patel · 2016
Later among the works it cites.
Information geometric approach to recursive update in nonlinear filtering
Yubo Li, Yongqiang Cheng, Xiang Li, Xiaoqiang Hua, and Yuliang Qin · 2017
Closest in time.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Yann Ollivier, Ludovic Arnold, Anne Auger, and Nikolaus Hansen · 2017
Closest in time.
Unbiased online recurrent optimization
Corentin Tallec and Yann Ollivier · 2017
Closest in time.