Fetching the paper…
Reading the bibliography…
Self-attention networks (SAN) have been introduced into automatic speech recognition (ASR) and achieved state-of-the-art performance owing to its superior ability in capturing long term dependency.
K. Chen and Q. Huo, · 1910
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Learning feature mapping using deep neural network bottleneck features for distant large vocabulary speech recognition,”
I. Himawan, P. Motlicek, D. Imseng, B. Potard, N. Kim, and J. Lee, · 2015
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Syllable-based acoustic modeling with ctc-smbr-lstm,”
Zhongdi Qu, Parisa Haghani, Eugene Weinstein, and Pedro Moreno, · 2017
Earlier work this paper cites.
“A time-restricted self-attention layer for asr,”
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Self-attentional acoustic models,”
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, · 2018
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Deep-fsmn for large vocabulary continuous speech recognition,”
Shiliang Zhang, Ming Lei, Zhijie Yan, and Lirong Dai, · 2018
Later among the works it cites.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Closest in time.
“Augmenting self-attention with persistent memory,”
S. Sainbayar, G. Edouard, L. Guillaume, J. Herve, and J. Armand, · 2019
Closest in time.
“Augmenting self-attention with persistent memory,”
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Hervé Jégou, and Armand Joulin, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.