Fetching the paper…
Reading the bibliography…
The autoregressive (AR) models, such as attention-based encoder-decoder models and RNN-Transducer, have achieved great success in speech recognition.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
C. J. K, B. Dzmitry, S. Dmitriy, C. Kyunghyun, and B. Yoshua, “Attention-based models for speech recognition,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
K. Suyoun, H. Takaaki, and W. Shinji, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in
2017
Earlier work this paper cites.
D. Linhao, X. Shuang, and X. Bo, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in
2018
Earlier work this paper cites.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang
2019
Earlier work this paper cites.
Z. Tian, J. Yi, J. Tao, Y. Bai, and Z. Wen, “Self-Attention Transducers for End-to-End Speech Recognition,” in
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in
2019
Cited alongside, same era.
Y. zhao, J. Li, X. Wang, and Y. Li, “The speechtransformer for large-scale mandarin chinese speech recognition,” in
2020
Later among the works it cites.
Y. Bai, J. Yi, J. Tao, Z. Tian, Z. Wen, and S. Zhang, “Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition,” in
2020
Later among the works it cites.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Deliberation model based two-pass end-to-end speech recognition,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Yu, W. Han, A. Gulati, C.-C. Chiu, B. Li, T. N. Sainath, Y. Wu, and R. Pang, “Dual-mode asr: Unify and improve streaming asr with full-context modeling,”
2021
Closest in time.
Z. Tian, J. Yi, J. Tao, Y. Bai, S. Zhang, and Z. Wen, “Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition,” in
2086
Closest in time.