Fetching the paper…
Reading the bibliography…
In this paper, we aim at providing an introduction to the gradient descent based optimization algorithms for learning deep neural network models.
A method of solving a convex programming problem with convergence rate O(1/sqr(k))
Yurii Nesterov · 1983
Earlier work this paper cites.
Convergence models of genetic algorithm selection schemes
Dirk Thierens and David E. Goldberg · 1994
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Genetic algorithms for evolving deep neural networks
Eli David and Iddo Greental · 2017
Cited alongside, same era.
Incorporating Nesterov Momentum into Adam
Timothy Dozat
Cited in the paper.
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2017
Later among the works it cites.
GADAM: genetic-evolutionary ADAM for deep neural network optimization
Jiawei Zhang and Fisher B. Gouza · 2018
Later among the works it cites.
SEGEN: sample-ensemble genetic evolutional network model
Jiawei Zhang and Fisher B. Gouza · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…