Fetching the paper…
Reading the bibliography…
The success of self-attention in NLP has led to recent applications in end-to-end encoder-decoder architectures for speech recognition.
“An application of recurrent nets to phone probability estimation,”
A.J. Robinson, · 1994
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Offline handwriting recognition with multidimensional recurrent neural networks,”
A. Graves and J. Schmidhuber, · 2009
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs,”
A.Y. Hannun, A.L. Maas, D. Jurafsky, and A.Y. Ng, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
I. Sutskever, O. Vinyals, and Q.V. Le, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Earlier work this paper cites.
“Deep Speech: Scaling up end-to-end speech recognition,”
A.Y. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al., · 2014
Earlier work this paper cites.
“Lexicon-free conversational speech recognition with neural networks,”
A.L. Maas, Z. Xie, D. Jurafsky, and A.Y. Ng, · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Earlier work this paper cites.
“Convolutional, long short-term memory, fully connected deep neural networks,”
T.N. Sainath, O. Vinyals, A. Senior, and H. Sak, · 2015
Earlier work this paper cites.
“MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems,”
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Connectionist temporal modeling for weakly supervised action labeling,”
D.-A. Huang, L. Fei-Fei, and J.C. Niebles, · 2016
Earlier work this paper cites.
“Deep Speech 2: End-to-end speech recognition in English and Mandarin,”
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al., · 2016
Cited alongside, same era.
“Wav2Letter: an end-to-end ConvNet-based speech recognition system,”
R. Collobert, C. Puhrsch, and G. Synnaeve, · 2016
Cited alongside, same era.
“Towards end-to-end speech recognition with deep convolutional neural networks,”
Y. Zhang, M. Pezeshki, P. Brakel, S. Zhang, C. Laurent, Y. Bengio, and A. C. Courville, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q.V. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Long short-term memory-networks for machine reading,”
J. Cheng, L. Dong, and M. Lapata, · 2016
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C.-C. Chiu, T.N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R.J. Weiss, K. Rao, K. Gonina, et al., · 2018
Later among the works it cites.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Later among the works it cites.
“Improved training of end-to-end attention models for speech recognition,”
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, · 2018
Later among the works it cites.
“DiSAN: Directional self-attention network for RNN/CNN-free language understanding,”
T. Shen, T. Zhou, G. Long, J. Jiang, S. Pan, and C. Zhang, · 2018
Later among the works it cites.
“A time-restricted self-attention layer for ASR,”
D. Povey, H. Hadian, P. Ghahremani, K. Li, and S. Khudanpur, · 2018
Later among the works it cites.
“Self-attentional acoustic models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Fully convolutional recurrent network for handwritten Chinese text recognition,”
Z. Xie, Z. Sun, L. Jin, Z. Feng, and S. Zhang, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Cited alongside, same era.
“Residual convolutional CTC networks for automatic speech recognition,”
Y. Wang, X. Deng, S. Pu, and Z. Huang, · 2017
Cited alongside, same era.
“Very deep convolutional networks for end-to-end speech recognition,”
Y. Zhang, W. Chan, and N. Jaitly, · 2017
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Cited alongside, same era.
“Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,”
T. Hori, S. Watanabe, Y. Zhang, and W. Chan, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, · 2018
Later among the works it cites.
“Syllable-based sequence-to-sequence speech recognition with the Transformer in Mandarin Chinese,”
S. Zhou, L. Dong, S. Xu, and B. Xu, · 2018
Later among the works it cites.
“Multilingual end-to-end speech recognition with a single Transformer on low-resource languages,”
S. Zhou, S. Xu, and B. Xu, · 2018
Later among the works it cites.
“Advancing connectionist temporal classification with attention modeling,”
Amit Das, Jinyu Li, Rui Zhao, and Yifan Gong, · 2018
Later among the works it cites.
“Improving end-to-end speech recognition with policy learning,”
Y. Zhou, C. Xiong, and R. Socher, · 2018
Later among the works it cites.
“Subword and crossword units for CTC acoustic models,”
T. Zenkel, R. Sanabria, F. Metze, and A. Waibel, · 2018
Later among the works it cites.
“End-to-end speech recognition from the raw waveform,”
N. Zeghidour, N. Usunier, G. Synnaeve, R. Collobert, and E. Dupoux, · 2018
Later among the works it cites.
“Optimal completion distillation for sequence learning,”
S. Sabour, W. Chan, and M. Norouzi, · 2018
Later among the works it cites.
“Learning noise-invariant representations for robust speech recognition,”
D. Liang, Z. Huang, and Z.C. Lipton, · 2018
Later among the works it cites.
“Cold Fusion: Training seq2seq models together with language models,”
A. Sriram, H. Jun, S. Satheesh, and A. Coates, · 2018
Later among the works it cites.