Fetching the paper…
Reading the bibliography…
End-to-end speech recognition has become popular in recent years, since it can integrate the acoustic, pronunciation and language models into a single neural network.
Y. Bengio, P. Simard, P. Frasconi
1994
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2011
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,”
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,”
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in
2017
Cited alongside, same era.
2018
Later among the works it cites.
S. Zhang, M. Lei, Y. Liu, and W. Li, “Investigation of modeling units for mandarin speech recognition using DFSMN-CTC-sMBR,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Tian, J. Yi, J. Tao, Y. Bai, and Z. Wen, “Self-attention transducers for end-to-end speech recognition,”
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Zhang, L. Ming, Z. Yan, and L. Dai, “Deep-FSMN for large vocabulary continuous speech recognition,” pp. 5869–5873, 2018
2018
Cited alongside, same era.
S. Zhang and M. Lei, “Acoustic modeling with DFSMN-CTC and joint CTC-CE learning.” in
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
C. Shan, C. Weng, G. Wang, D. Su, M. Luo, D. Yu, and L. Xie, “Component fusion: Learning replaceable language model component for end-to-end speech recognition system,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.