P. Potash, A. Romanov, and A. Rumshisky, “Ghostwriter: using an lstm for automatic rap lyric generation,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , 2015, pp. 1919– 1924
1924
Earlier work this paper cites.
L. Theis and M. Bethge, “Generative image modeling using spatial lstms,” in Advances in Neural Information Processing Systems , 2015, pp. 1918–1926
1926
Earlier work this paper cites.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” IEEE transactions on acoustics, speech, and signal processing , vol. 37, no. 3, pp. 328–339, 1989
1989
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, 1990
1990
Earlier work this paper cites.
R. J. Williams, “Training recurrent networks using the extended kalman filter,” in Neural Networks, 1992. IJCNN., International Joint Conference on , vol. 4. IEEE, 1992, pp. 241–246
1992
Earlier work this paper cites.
C. Tuerk and T. Robinson, “Speech synthesis using artificial neural networks trained on cepstral coefficients.” in EUROSPEECH , 1993
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE transactions on neural networks , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Haykin, Neural networks: a comprehensive foundation . Prentice Hall PTR, 1994
1994
Earlier work this paper cites.
G. V. Puskorius and L. A. Feldkamp, “Neurocontrol of nonlinear dynamical systems with kalman filter trained recurrent networks,” IEEE Transactions on neural networks , vol. 5, no. 2, pp. 279–297, 1994
1994
Earlier work this paper cites.
P. J. Angeline, G. M. Saunders, and J. B. Pollack, “An evolutionary algorithm that constructs recurrent neural networks,” IEEE transactions on Neural Networks , vol. 5, no. 1, pp. 54–65, 1994
1994
Earlier work this paper cites.
K. Unnikrishnan and K. P. Venugopal, “Alopex: A correlation-based learning algorithm for feedforward and recurrent neural networks,” Neural Computation , vol. 6, no. 3, pp. 469–490, 1994
1994
Earlier work this paper cites.
R. J. Williams and D. Zipser, “Gradient-based learning algorithms for recurrent networks and their computational complexity,” Backpropagation: Theory, architectures, and applications , vol. 1, pp. 433–486, 1995
1995
Earlier work this paper cites.
M. Schuster, “Bi-directional recurrent neural networks for speech recognition,” Technical report, Tech. Rep., 1996
1996
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
S. Ma and C. Ji, “A unified approach on fast training of feedforward and recurrent networks using em algorithm,” IEEE transactions on signal processing , vol. 46, no. 8, pp. 2270–2274, 1998
1998
Earlier work this paper cites.
O. Karaali, G. Corrigan, I. Gerson, and N. Massey, “Text-to-speech conversion with neural networks: A recurrent tdnn approach,” arXiv preprint cs/9811032 , 1998
1998
Earlier work this paper cites.
L.-W. Chan and C.-C. Szeto, “Training recurrent network with block-diagonal approximated levenberg-marquardt algorithm,” in Neural Networks, 1999. IJCNN’99. International Joint Conference on , vol. 3. IEEE, 1999, pp. 1521–1526
1999
Earlier work this paper cites.
F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget: Continual prediction with lstm,” 1999
1999
Earlier work this paper cites.
S. S. Haykin et al. , Kalman filtering and neural networks . Wiley Online Library, 2001
2001
Earlier work this paper cites.
D. Eck and J. Schmidhuber, “A first look at music composition using lstm recurrent neural networks,” Istituto Dalle Molle Di Studi Sull Intelligenza Artificiale , vol. 103, 2002
2002
Earlier work this paper cites.
——, “Finding temporal structure in music: Blues improvisation with lstm recurrent networks,” in Neural Networks for Signal Processing, 2002. Proceedings of the 2002 12th IEEE Workshop on . IEEE, 2002, pp. 747–756
2002
Earlier work this paper cites.
J. A. Pérez-Ortiz, F. A. Gers, D. Eck, and J. Schmidhuber, “Kalman filters improve lstm network performance in problems unsolvable by traditional recurrent nets,” Neural Networks , vol. 16, no. 2, pp. 241–250, 2003
2003
Earlier work this paper cites.
P. Baldi and G. Pollastri, “The principled design of large-scale recursive neural network architectures–dag-rnns and the protein structure prediction problem,” Journal of Machine Learning Research , vol. 4, no. Sep, pp. 575–602, 2003
2003
Earlier work this paper cites.
L. Bottou, “Stochastic learning,” in Advanced lectures on machine learning . Springer, 2004, pp. 146–168
2004
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks , vol. 18, no. 5, pp. 602–610, 2005
2005
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Earlier work this paper cites.
C. M. Bishop, Pattern recognition and machine learning . springer, 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
Y. Bengio, Y. LeCun et al. , “Scaling learning algorithms towards ai,” Large-scale kernel machines , vol. 34, no. 5, pp. 1–41, 2007
2007
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in Advances in neural information processing systems , 2007, pp. 153–160
2007
Earlier work this paper cites.
A. Graves, S. Fernandez, and J. Schmidhuber, “Multi-dimensional recurrent neural networks,” arXiv preprint arXiv:0705.2011v1 , 2007
Original
2007
Earlier work this paper cites.
M. Wöllmer, F. Eyben, S. Reiter, B. Schuller, C. Cox, E. Douglas-Cowie, and R. Cowie, “Abandoning emotion classes-towards continuous emotion recognition with modelling of long-range dependencies,” in Ninth Annual Conference of the International Speech Communication Association , 2008
2008
Earlier work this paper cites.
A. Graves, M. Liwicki, H. Bunke, J. Schmidhuber, and S. Fernández, “Unconstrained on-line handwriting recognition with recurrent neural networks,” in Advances in Neural Information Processing Systems , 2008, pp. 577–584
2008
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Offline handwriting recognition with multidimensional recurrent neural networks,” in Advances in neural information processing systems , 2009, pp. 545–552
2009
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, “Recurrent neural network based language model.” in Interspeech , vol. 2, 2010, p. 3
2010
Earlier work this paper cites.
J. Martens, “Deep learning via hessian-free optimization,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10) , 2010, pp. 735–742
2010
Earlier work this paper cites.
D. T. Mirikitani and N. Nikolaev, “Recursive bayesian recurrent neural networks for time-series modeling,” IEEE Transactions on Neural Networks , vol. 21, no. 2, pp. 262–274, 2010
2010
Earlier work this paper cites.
I. Sutskever, J. Martens, and G. E. Hinton, “Generating text with recurrent neural networks,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 1017–1024
2011
Earlier work this paper cites.
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan, “Better mini-batch algorithms via accelerated gradient methods,” in Advances in neural information processing systems , 2011, pp. 1647–1655
2011
Earlier work this paper cites.
J. Martens and I. Sutskever, “Learning recurrent neural networks with hessian-free optimization,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 1033–1040
2011
Earlier work this paper cites.
A. Graves, S. Fernández, and J. Schmidhuber, “Multi-dimensional recurrent neural networks,” 2007. [Online]. Available: http://arxiv.org/abs/0705.2011
Original
2011
Earlier work this paper cites.