Fetching the paper…
Reading the bibliography…
Recently, Transformer has gained success in automatic speech recognition (ASR) field.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Hkust/mts: A very large scale mandarin telephone speech corpus,”
Y. Liu, P Fung, Y. Yang, C Cieri, S. Huang, and D. Graff, · 2006
Earlier work this paper cites.
Supervised sequence labelling with recurrent neural networks
K. Kawakami, · 2008
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D., K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al., · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
L. J. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Cited alongside, same era.
“Rethinking the inception architecture for computer vision,”
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for asr based on lattice-free mmi,”
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, · 2016
Cited alongside, same era.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Cited alongside, same era.
“Joint CTC/attention decoding for end-to-end speech recognition,”
T. Hori, S. Watanabe, and J. Hershey, · 2017
Cited alongside, same era.
“Attention is all you need,”
“Espnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, and N. Chen, · 2018
Later among the works it cites.
“Online Hybrid CTC/Attention Architecture for End-to-End Speech Recognition,”
H. Miao, G. Cheng, P. Zhang, L. Ta, and Y. Yan, · 2019
Later among the works it cites.
“Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration,”
S. Karita, N. E. Y. Soplin, S. Watanabe, M. Delcroix, A. Ogawa, and T. Nakatani, · 2019
Later among the works it cites.
“Very Deep Self-Attention Networks for End-to-End Speech Recognition,”
N. Pham, T. Nguyen, J., M. Müller, and A. Waibel, · 2019
Later among the works it cites.
“A comparative study on transformer vs rnn in speech applications,”
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. Yalta, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, and W. Zhang, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, · 2018
Cited alongside, same era.
“Speech-transformer: A norecurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
C. Chiu and C. Raffel, · 2018
Cited alongside, same era.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. V. Le, and R. Salakhutdinov, · 2019
Later among the works it cites.
L. Dong, F. Wang, and B. Xu, · 2019
Later among the works it cites.
“Streaming attention,” https://github.com/HaoranMiao/streaming-attention, 2020,
H. Miao and G. Cheng, · 2020
Closest in time.