Fetching the paper…
Reading the bibliography…
The RNN-Transducers and improved attention-based encoder-decoder models are widely applied to streaming speech recognition.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
2010
Earlier work this paper cites.
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE international conference on acoustics, speech and signal processing . IEEE, 2013, pp. 6645–6649
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 167–174
2015
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International conference on machine learning , 2016, pp. 173–182
2016
Earlier work this paper cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi.” in Interspeech , 2016, pp. 2751–2755
2016
Earlier work this paper cites.
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,” arXiv preprint arXiv:1712.05382 , 2017
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid ctc/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
T. N. Sainath, C.-C. Chiu, R. Prabhavalkar, A. Kannan, Y. Wu, P. Nguyen, and Z. Chen, “Improving the performance of online neural transducer models,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5864–5868
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. He, T. N. Sainath, R. Prabhavalkar, I. Mcgraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices.” arXiv: Computation and Language , 2018
2018
Cited alongside, same era.
T. Bagby, K. Rao, and K. C. Sim, “Efficient implementation of recurrent neural network transducer in tensorflow,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 506–512
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Cited alongside, same era.
N. Moritz, T. Hori, and J. Le Roux, “Triggered attention for end-to-end speech recognition,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5666–5670
2019
Cited alongside, same era.
Z. Tian, J. Yi, J. Tao, Y. Bai, and Z. Wen, “Self-Attention Transducers for End-to-End Speech Recognition,” in Proc. Interspeech 2019 , 2019, pp. 4395–4399. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2203
2019
Cited alongside, same era.
J. Li, R. Zhao, H. Hu, and Y. Gong, “Improving rnn transducer modeling for end-to-end speech recognition,” arXiv: Computation and Language , 2019
2019
Cited alongside, same era.
L. Dong and B. Xu, “Cif: Continuous integrate-and-fire for end-to-end speech recognition,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6079–6083
2020
Closest in time.
Z. Tian, J. Yi, Y. Bai, J. Tao, S. Zhang, and Z. Wen, “Synchronous transformers for end-to-end speech recognition,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7884–7888
2020
Closest in time.
2020
Closest in time.
W. Zhou, S. Ralf, and N. Hermann, “Robust beam search for encoder-decoder attention based speech recognition without length bias,” in Proc. Interspeech 2020 , 2020, pp. 1768–1772
2020
Closest in time.