Fetching the paper…
Reading the bibliography…
Gradient descent algorithms have been used in countless applications since the inception of Newton's method.
H. Robbins, S. Monro, A Stochastic Approximation Method, Annals of Mathematical Statistics 22.3 (1951), 400-407
1951
Earlier work this paper cites.
D.E. Rumelhart, G. Hinton, R.J. Williams, Learning representations by back-propagating errors, Nature 323 (1986), 533-536
1986
Earlier work this paper cites.
G. Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signal Systems 2 (1989), 303-314
1989
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition, Proceedings of the IEEE (1998)
1998
Earlier work this paper cites.
J. Hubbard, D. Schleicher, S. Sutherland, How to find all roots of complex polynomials by Newton’s method, Inventiones mathematicae 146 (2001), 1-33
2001
Earlier work this paper cites.
A. Krizhevsky, G. Hinton, Learning Multiple Layers of Features from Tiny Images, Technical Report, University of Toronto (2009)
2009
Earlier work this paper cites.
J. Duchi, E. Hazan, Y. Singer, Adaptive Subgradient Methods for Online Learning and Stochastic Optimization, Journal of Machine Learning Research 12 (2011), 2121-2159
2011
Cited alongside, same era.
D.P. Kingma, J. Ba, Adam: A Method for Stochastic Optimization, arXiv preprint
2014
Cited alongside, same era.
K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, arXiv preprint
2015
Cited alongside, same era.
M.A. Nielsen, Neural Networks and Deep Learning, Determination Press (2015)
2015
Cited alongside, same era.
K. Simonyan, A. Zisserman, Very Deep Convolutional Networks for Large-Scale Image Recognition, ICLR (2015)
2015
Cited alongside, same era.
T. Dozat, Incorporating Nesterov Momentum into Adam, ICLR Workshop 1 (2016), 2013-2016
I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press (2016), www.deeplearningbook.org
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, J. Sun, Identity Mappings in Deep Residual Networks, arXiv preprint
2016
Later among the works it cites.
S. Ruder, An overview of gradient descent optimization algorithms, arXiv preprint
2016
Later among the works it cites.
2016
Later among the works it cites.
R. Koss, Complex Varieties as Minima, MA Thesis, Eastern Illinois University (2019)
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
G. Hinton, Neural Networks for Machine Learning, Lecture 6e rmsprop: Divide the gradient by a running average of its recent magnitude, www.cs.toronto.edu/ tijmen/csc321/slides/lecture_slides_lec6.pdf
Cited in the paper.