Fetching the paper…
Reading the bibliography…
We present a novel per-dimension learning rate method for gradient descent called ADADELTA.
“A stochastic approximation method,”
H. Robinds and S. Monro, · 1951
Earlier work this paper cites.
“Learning representations by back-propagating errors,”
D.E. Rumelhart, G.E. Hinton, and R.J. Williams, · 1986
Earlier work this paper cites.
“Improving the convergence of back-propagation learning with second order methods,”
S. Becker and Y. LeCun, · 1988
Earlier work this paper cites.
“Adaptive subgradient methods for online leaning and stochastic optimization,”
J. Duchi, E. Hazan, and Y. Singer, · 2010
Cited alongside, same era.
“Large scale distributed deep networks,”
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng, · 2012
Cited alongside, same era.
“No more pesky learning rates,” arXiv:1206.1106, 2012
T. Schaul, S. Zhang, and Y. LeCun, · 2012
Closest in time.
“Application of pretrained deep neural networks to large vocabulary speech recognition,”
N. Jaitly, P. Nguyen, A. Senior, and V. Vanhoucke, · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…