Fetching the paper…
Reading the bibliography…
We study two types of preconditioners and preconditioned stochastic gradient descent (SGD) methods in a unified framework.
A method of solving a convex programming problem with convergence rate o(1/srt(k))
Y. Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Equivariant adaptive source separation
J. F. Cardoso and B. H. Laheld · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
D. John, H. Elad, and S. Yoram · 2011
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
K. Alex, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization
Y. N. Dauphin, H. Vries, and Y. Bengio · 2015
Cited alongside, same era.
Adam: a method for stochastic optimization
D. P. Kingma and J. L. Ba · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. B. Grosse · 2015
Later among the works it cites.
Parallel training of DNNs with natural gradient and parameter averaging
D. Povey, X. Zhang, and S. Khudanpur · 2015
Later among the works it cites.
Preconditioned stochastic gradient descent
X. L. Li · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…