Fetching the paper…
Reading the bibliography…
End-to-end models are favored in automatic speech recognition (ASR) because of their simplified system structure and superior performance.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The KALDI speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” Computer Science , 2012
2012
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in ICASSP , 2013, pp. 6645–6649
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in ICML , 2014, pp. 1764–1772
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP , 2016, pp. 4960–4964
2016
Earlier work this paper cites.
E. Battenberg, J. Chen, R. Child, A. Coates, Y. G. Y. Li, H. Liu, S. Satheesh, A. Sriram, and Z. Zhu, “Exploring neural transducers for end-to-end speech recognition,” in ASRU , 2017, pp. 206–213
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in COCOSDA , 2017, pp. 1–5
2017
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in ICASSP , 2018, pp. 5884–5888
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPNet: End-to-end speech processing toolkit,” INTERSPEECH , 2018
2018
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in INTERSPEECH , 2020, pp. 5036–5040
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Wang, Z. Yao, X. Shi, and L. Xie, “Cascade RNN-transducer: Syllable based streaming on-device Mandarin speech recognition with a syllable-to-character converter,” in SLT , 2021, pp. 15–21
2021
Closest in time.
Y. Zhang, S. Sun, and L. Ma, “Tiny transducer: A highly-efficient speech recognition model on edge devices,” ICASSP , 2021
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K. Chai Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein, “Streaming end-to-end speech recognition for mobile devices,” in ICASSP , 2019, pp. 6381–6385
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” INTERSPEECH , pp. 2613–2617, 2019
2019
Cited alongside, same era.
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, “Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,” ICASSP , 2021
2021
Closest in time.
H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang, “Informer: Beyond efficient transformer for long sequence time-series forecasting,” AAAI , 2021
2021
Closest in time.