Fetching the paper…
Reading the bibliography…
We present a unifying framework for adapting the update direction in gradient-based iterative optimization methods.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R Salakhutdinov · 2006
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. J. Wright · 2006
Cited alongside, same era.
Essential Mathematics for Economic Analysis
K. Sydsaeter and P. Hammond · 2008
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
P-Y. Massé and Y. Ollivier · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…