Fetching the paper…
Reading the bibliography…
Deep neural network learning can be formulated as a non-convex optimization problem.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Genetic algorithms
J. Holland · 1992
Earlier work this paper cites.
Parameterisation of a stochastic model for human face identification
F. Samaria and A. C. Harter · 1994
Earlier work this paper cites.
Convergence models of genetic algorithm selection schemes
D. Thierens and D. Goldberg · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
An Introduction to Genetic Algorithms
M. Mitchell · 1998
Earlier work this paper cites.
Training invariant support vector machines
D. Decoste and B. Schölkopf · 2002
Earlier work this paper cites.
Tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the “echo state network” approach
H. Jaeger · 2002
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. Hinton, S. Osindero, and Y. Teh · 2006
Earlier work this paper cites.
A scalable hierarchical distributed language model
A. Mnih and G. Hinton · 2009
Earlier work this paper cites.
Semantic hashing
R. Salakhutdinov and G. Hinton · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P. Manzagol · 2010
Cited alongside, same era.
Large scale image annotation: Learning to rank with joint word-image embeddings
J. Weston, S. Bengio, and N. Usunier · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Wsabie: Scaling up to large vocabulary image annotation
J. Weston, S. Bengio, and N. Usunier · 2011
Cited alongside, same era.
Deep neural network language models
E. Arisoy, T. Sainath, B. Kingsbury, and B. Ramabhadran · 2012
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
Genetic algorithms for evolving deep neural networks
O. E. D. and I. G · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Later among the works it cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Later among the works it cites.
Incorporating nesterov momentum into adam
T. Dozat · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A practical guide to training restricted boltzmann machines
G. Hinton · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
New types of deep neural network learning for speech recognition and related applications: An overview
L. Deng, G. Hinton, and B. Kingsbury · 2013
Cited alongside, same era.
E. David and I. Greental · 2017
Later among the works it cites.
Improving generalization performance by switching from adam to SGD
N. Keskar and R. Socher · 2017
Later among the works it cites.
F. Such, V. Madhavan, E. Conti, J. Lehman, K. Stanley, and J. Clune · 2017
Later among the works it cites.
Deep forest: Towards an alternative to deep neural networks
Z. Zhou and J. Feng · 2017
Later among the works it cites.
segen: Sample-ensemble genetic evolutional network model
J. Zhang, L. Cui, and F. Gouza · 2018
Closest in time.