Fetching the paper…
Reading the bibliography…
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC).
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
Matrix equation X A + B X = C XA+BX=C
R. Smith · 1968
Earlier work this paper cites.
The Levenberg-Marquardt algorithm: implementation and theory
J. Moré · 1978
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k ) \mathcal{O}(1/\sqrt{k})
Y. Nesterov · 1983
Earlier work this paper cites.
Experiments on learning by back propagation
D. Plaut, S. Nowlan, and G. E. Hinton · 1986
Earlier work this paper cites.
The solution of the matrix equations A X B − C X D = E AXB-CXD=E and ( Y A − D Z , Y C − B Z ) = ( E , F ) (YA-DZ,YC-BZ)=(E,F)
K.-w. E. Chu · 1987
Earlier work this paper cites.
Improving the Convergence of Back-Propagation Learning with Second Order Methods
S. Becker and Y. LeCun · 1989
Earlier work this paper cites.
Note on learning rate schedules for stochastic optimization
C. Darken and J. E. Moody · 1990
Earlier work this paper cites.
An improved newton iteration for the generalized inverse of a matrix, with applications
V. Pan and R. Schreiber · 1991
Earlier work this paper cites.
Solution of the sylvester matrix equation A X B T + C X D T = E AXB^{T}+CXD^{T}=E
J. D. Gardiner, A. J. Laub, J. J. Amato, and C. B. Moler · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Müller · 1998
Earlier work this paper cites.
Centering neural network gradient factors
N. N. Schraudolph · 1998
Earlier work this paper cites.
Joint mean-covariance models with applications to longitudinal data: unconstrained parameterisation
M. Pourahmadi · 1999
Earlier work this paper cites.
Matrix momentum for practical natural gradient learning
S. Scarpetta, M. Rattray, and D. Saad · 1999
Earlier work this paper cites.
Methods of Information Geometry , volume 191 of Translations of Mathematical monographs
S.-I. Amari and H. Nagaoka · 2000
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
T. Heskes · 2000
Earlier work this paper cites.
Adaptive natural gradient learning algorithms for various stochastic models
H. Park, S.-I. Amari, and K. Fukumizu · 2000
Cited alongside, same era.
The ubiquitous kronecker product
C. F. Van Loan · 2000
Cited alongside, same era.
Fast curvature matrix-vector products for second-order gradient descent
N. N. Schraudolph · 2002
Cited alongside, same era.
Sharpness in rates of convergence for CG and symmetric Lanczos methods
R.-C. Li · 2005
Cited alongside, same era.
Pattern Recognition and Machine Learning (Information Science and Statistics)
C. M. Bishop · 2006
Cited alongside, same era.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Cited alongside, same era.
Training deep and recurrent networks with Hessian-free optimization
J. Martens and I. Sutskever · 2012
Later among the works it cites.
Estimating the Hessian by backpropagating curvature
J. Martens, I. Sutskever, and K. Swersky · 2012
Later among the works it cites.
Deep learning made easier by linear transformations in perceptrons
T. Raiko, H. Valpola, and Y. LeCun · 2012
Later among the works it cites.
Krylov subspace descent for deep learning
O. Vinyals and D. Povey · 2012
Later among the works it cites.
Training neural networks with stochastic Hessian-free optimization
R. Kiros · 2013
Later among the works it cites.
Riemannian metrics for neural networks
Y. Ollivier · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Numerical optimization
J. Nocedal and S. J. Wright · 2006
Cited alongside, same era.
A stochastic quasi-newton method for online convex optimization
N. N. Schraudolph, J. Yu, and S. G�nter · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.-a. Manzagol, and Y. Bengio · 2008
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
A tutorial on stochastic approximation algorithms for training restricted boltzmann machines and deep belief nets
K. Swersky, B. Chen, B. Marlin, and N. de Freitas · 2010
Cited alongside, same era.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
T. Vatanen, T. Raiko, H. Valpola, and Y. LeCun · 2013
Later among the works it cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2013
Later among the works it cites.
New insights and perspectives on the natural gradient method
J. Martens · 2014
Later among the works it cites.
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Later among the works it cites.
Computational methods for linear matrix equations
V. Simoncini · 2014
Later among the works it cites.
Mean-normalized stochastic gradient for large-scale deep learning
S. Wiesler, A. Richard, R. Schlüter, and H. Ney · 2014
Later among the works it cites.
Scaling up natural gradient by factorizing fisher information
R. Grosse and R. Salakhutdinov · 2015
Closest in time.
Parallel training of DNNs with natural gradient and parameter averaging
D. Povey, X. Zhang, and S. Khudanpur · 2015
Closest in time.