“Some methods of speeding up the convergence of iteration methods,”
Boris T Polyak, · 1964
Earlier work this paper cites.
“An application of recurrent nets to phone probability estimation,”
Anthony J Robinson, · 1994
Earlier work this paper cites.
Connectionist speech recognition: a hybrid approach
Hervé Bourlard and Nelson Morgan, · 1994
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K Paliwal, · 1997
Earlier work this paper cites.
“Learning precise timing with LSTM recurrent networks,”
Felix A Gers, Nicol N Schraudolph, and Jürgen Schmidhuber, · 2003
Earlier work this paper cites.
“Framewise phoneme classification with bidirectional lstm and other neural network architectures,”
Alex Graves and Jürgen Schmidhuber, · 2005
Earlier work this paper cites.
“Gammatone features and feature combination for large vocabulary speech recognition,”
Ralf Schlüter, L Bezrukov, Hannes Wagner, and Hermann Ney, · 2007
Earlier work this paper cites.
“The RWTH 2009 Quaero ASR evaluation system for English and German,”
Markus Nußbaum-Thom, Simon Wiesler, Martin Sundermeyer, Christian Plahl, Stefan Hahn, Ralf Schlüter, and Hermann Ney, · 2010
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“RASR - the RWTH Aachen university open source speech recognition toolkit,”
David Rybach, Stefan Hahn, Patrick Lehnen, David Nolden, Martin Sundermeyer, Zoltan Tüske, Simon Wiesler, Ralf Schlüter, and Hermann Ney, · 2011
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2011
Earlier work this paper cites.
“Feature engineering in context-dependent deep neural networks for conversational speech transcription,”
Frank Seide, Gang Li, Xie Chen, and Dong Yu, · 2011
Earlier work this paper cites.
“Theano: new features and speed improvements,” Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop, 2012
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian J. Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio, · 2012
Earlier work this paper cites.
“Adadelta: An adaptive learning rate method,”
Original
Matthew D Zeiler, · 2012
Earlier work this paper cites.
“Lecture 6.5 - RMSprop: Divide the gradient by a running average of its recent magnitude.,” COURSERA: Neural Networks for Machine Learning, 2012
Tijmen Tieleman and Geoffrey Hinton, · 2012
Earlier work this paper cites.