Fetching the paper…
Reading the bibliography…
Many of the current state-of-the-art Large Vocabulary Continuous Speech Recognition Systems (LVCSR) are hybrids of neural networks and Hidden Markov Models (HMMs).
Review of neural networks for speech recognition
Lippmann, R. P. (1989) · 1989
Earlier work this paper cites.
The use of recurrent neural networks in continuous speech recognition
Robinson, T., Hochberg, M., and Renals, S. (1996) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
Mohri, M., Pereira, F., and Riley, M. (2002) · 2002
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J. (2006) · 2006
Earlier work this paper cites.
OpenFst: A general and efficient weighted finite-state transducer library
Allauzen, C., Riley, M., Schalkwyk, J., Skut, W., and Mohri, M. (2007) · 2007
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A. (2012) · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Zeiler, M. D. (2012) · 2012
Cited alongside, same era.
High-dimensional sequence transduction
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P. (2013) · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Graves, A. (2013) · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Closest in time.
Chan, W., Jaitly, N., Le, Q. V., and Vinyals, O. (2015) · 2015
Closest in time.
Attention-based models for speech recognition
Chorowski, J., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y. (2015) · 2015
Closest in time.
Gated feedback recurrent neural networks
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2015) · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chorowski, J., Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A. and Jaitly, N. (2014) · 2014
Cited alongside, same era.
A clockwork RNN
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J. (2014) · 2014
Cited alongside, same era.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A., et al. (2014) · 2014
Cited alongside, same era.
Deepspeech: Scaling up end-to-end speech recognition
Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., Coates, A., et al. (2014a)
Cited in the paper.
First-pass large vocabulary continuous speech recognition using bi-directional recurrent dnns
Hannun, A. Y., Maas, A. L., Jurafsky, D., and Ng, A. Y. (2014b)
Cited in the paper.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., and Kingsbury, B. (2012a)
Cited in the paper.
Gulcehre, C., Firat, O., Xu, K., Cho, K., Barrault, L., Lin, H.-C., Bougares, F., Schwenk, H., and Bengio, Y. (2015) · 2015
Closest in time.
EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding
Miao, Y., Gowayyed, M., and Metze, F. (2015) · 2015
Closest in time.
Blocks and fuel: Frameworks for deep learning
van Merriënboer, B., Bahdanau, D., Dumoulin, V., Serdyuk, D., Warde-Farley, D., Chorowski, J., and Bengio, Y. (2015) · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015) · 2015
Closest in time.