Fetching the paper…
Reading the bibliography…
This paper studies the performance of a recently proposed preconditioned stochastic gradient descent (PSGD) algorithm on recurrent neural network (RNN) training.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proc. IEEE , vol. 78, no. 10, pp. 1550–1560, Oct. 1990
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no.8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
N. Schraudolph, “Fast curvature matrix-vector products for second-order gradient descent,” Neural Computation , vol. 14, no. 7, pp. 1723–1738, 2002
2002
Earlier work this paper cites.
J. Martens and I. Sutskever, “Learning recurrent neural networks with Hessian-free optimization,” In Proc. of the 28th ICML , 2011
2011
Cited alongside, same era.
2012
Cited alongside, same era.
Cited in the paper.
2015
Later among the works it cites.
X.-L. Li, “Preconditioned stochastic gradient descent,” arXiv:1512.04202, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…