Fetching the paper…
Reading the bibliography…
After a more than decade-long period of relatively little research activity in the area of recurrent neural networks, several new developments will be reviewed here that have allowed substantial progress both in understanding and in technical solutions towards more efficient training of recurrent networks.
“A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2}) ,”
Yu Nesterov, · 1983
Earlier work this paper cites.
“Learning representations by back-propagating errors,”
D.E. Rumelhart, G.E. Hinton, and R.J. Williams, · 1986
Earlier work this paper cites.
“ Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, T.U. Münich,” 1991
S. Hochreiter, · 1991
Earlier work this paper cites.
“Learning long-term dependencies with gradient descent is difficult,”
Y. Bengio, P. Simard, and P. Frasconi, · 1994
Earlier work this paper cites.
“Learning long-term dependencies is not as difficult with NARX recurrent neural networks,”
T. Lin, B. G. Horne, P. Tino, and C. L. Giles, · 1995
Earlier work this paper cites.
“Hierarchical recurrent neural networks for long-term dependencies,”
S. ElHihi and Y. Bengio, · 1996
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“A fast learning algorithm for deep belief nets,”
G. E. Hinton, S. Osindero, and Y.-W. Teh, · 2006
Earlier work this paper cites.
“Greedy layer-wise training of deep networks,”
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, · 2007
Earlier work this paper cites.
“Efficient learning of sparse representations with an energy-based model,”
M. Ranzato, C. Poultney, S. Chopra, and Y. LeCun, · 2007
Earlier work this paper cites.
“Optimization and applications of echo state networks with leaky- integrator neurons,”
Herbert Jaeger, Mantas Lukosevicius, Dan Popovici, and Udo Siewert, · 2007
Earlier work this paper cites.
“Echo-state networks with band-pass neurons: Towards generic time-scale-independent reservoir structures,”
Udo Siewert and Welf Wustlich, · 2007
Cited alongside, same era.
Learning deep architectures for AI
Yoshua Bengio, · 2009
Cited alongside, same era.
“The recurrent temporal restricted Boltzmann machine,”
I. Sutskever, G. Hinton, and G. Taylor, · 2009
Cited alongside, same era.
“Why does unsupervised pre-training help deep learning?,”
D. Erhan, Y. Bengio, A. Courville, P. Manzagol, P. Vincent, and S. Bengio, · 2010
Cited alongside, same era.
“Deep learning via Hessian-free optimization,”
J. Martens, · 2010
Cited alongside, same era.
“Temporal kernel recurrent neural networks,”
I. Sutskever and G. Hinton, · 2010
Cited alongside, same era.
“Extensions of recurrent neural network language model,”
Tomas Mikolov, Stefan Kombrink, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur, · 2011
Later among the works it cites.
Training Recurrent Neural Networks
I. Sutskever, · 2012
Closest in time.
“Unsupervised feature learning and deep learning: A review and new perspectives,”
Y. Bengio, A. Courville, and P. Vincent, · 2012
Closest in time.
“Building high-level features using large scale unsupervised learning,”
Q. Le, M. Ranzato, R. Monga, M. Devin, G. Corrado, K. Chen, J. Dean, and A. Ng, · 2012
Closest in time.
“Understanding the exploding gradient problem,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Rectified linear units improve restricted Boltzmann machines,”
V. Nair and G.E. Hinton, · 2010
Cited alongside, same era.
“Learning recurrent neural networks with Hessian-free optimization,”
J. Martens and I. Sutskever, · 2011
Cited alongside, same era.
“Extensions of recurrent neural network language model,”
T. Mikolov, S. Kombrink, L. Burget, J. Cernocky, and S. Khudanpur, · 2011
Cited alongside, same era.
“The Neural Autoregressive Distribution Estimator,”
H. Larochelle and I. Murray, · 2011
Cited alongside, same era.
“Deep sparse rectifier neural networks,”
X. Glorot, A. Bordes, and Y. Bengio, · 2011
Cited alongside, same era.
Statistical Language Models based on Neural Networks
Tomas Mikolov, · 2012
Closest in time.
“Context dependent reucrrent neural network language model,” Workshop on Spoken Language Technology, 2012
Tomas Mikolov and Geoffrey Zweig, · 2012
Closest in time.
“Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription,”
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, · 2012
Closest in time.
“ImageNet classification with deep convolutional neural networks,”
A. Krizhevsky, I. Sutskever, and G. Hinton, · 2012
Closest in time.
“Random search for hyper-parameter optimization,”
James Bergstra and Yoshua Bengio, · 2012
Closest in time.