Fetching the paper…
Reading the bibliography…
Encoder-decoder based sequence-to-sequence models have demonstrated state-of-the-art results in end-to-end automatic speech recognition (ASR).
“CSR-II (WSJ1) complete,” vol. LDC94S13A. Philadelphia: Linguistic Data Consortium, 1994
1994
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. International Conference on Machine Learning (ICML) , vol. 148, Jun. 2006, pp. 369–376
2006
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2015, pp. 5206–5210
2015
Earlier work this paper cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Proc. ISCA Interspeech , Sep. 2016, pp. 2751–2755
2016
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in Proc. ISCA Interspeech , Sep. 2017, pp. 939–943
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” J. Sel. Topics Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , Dec. 2017, pp. 6000–6010
2017
Cited alongside, same era.
T. Hori, S. Watanabe, Y. Zhang, and W. Chan, “Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,” in Proc. ISCA Interspeech , Aug. 2017, pp. 949–953
2017
Cited alongside, same era.
S. Karita, N. Yalta, S. Watanabe, M. Delcroix, A. Ogawa, and T. Nakatani, “Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration,” in Proc. ISCA Interspeech , Sep. 2019, pp. 1408–1412
2019
Later among the works it cites.
J. Schalkwyk, “An all-neural on-device speech recognizer,” Mar. 2019, url: https://ai.googleblog.com/2019/03/an-all-neural-on-device-speech.html
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Moritz, T. Hori, and J. Le Roux, “Triggered attention for end-to-end speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2019, pp. 5666–5670
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Chiu and C. Raffel, “Monotonic chunkwise attention,” in Proc. International Conference on Learning Representations (ICLR) , Apr. 2018
2018
Cited alongside, same era.
D. Povey, H. Hadian, P. Ghahremani, K. Li, and S. Khudanpur, “A time-restricted self-attention layer for ASR,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5874–5878
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5884–5888
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Moritz, T. Hori, and J. Le Roux, “Streaming end-to-end speech recognition with joint CTC-attention based models,” in Proc. IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , Dec. 2019, pp. 936–943
2019
Later among the works it cites.
N. Moritz, T. Hori, and J. Le Roux, “Unidirectional neural network architectures for end-to-end automatic speech recognition,” in Proc. ISCA Interspeech , Sep. 2019, pp. 76–80
2019
Later among the works it cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, and W. Zhang, “A comparative study on transformer vs RNN in speech applications,” in Proc. IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , Dec. 2019
2019
Later among the works it cites.