Fetching the paper…
Reading the bibliography…
The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks (RNNs) in end-to-end (E2E) automatic speech recognition (ASR) systems.
Connectionist Speech Recognition
H. Bourlard and N. Morgan, · 1994
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K Paliwal, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Hybrid speech recognition with deep bidirectional lstm,”
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Recurrent deep neural networks for robust speech recognition,”
C. Weng, D. Yu, S. Watanabe, and B. F. Juang, · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Acoustic modelling with CD-CTC-SMBR LSTM RNNs,”
Haşim Sak, Félix de Chaumont Quitry, Tara Sainath, Kanishka Rao, et al., · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K. Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
Navdeep Jaitly, David Sussillo, Quoc V Le, Oriol Vinyals, Ilya Sutskever, and Samy Bengio, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
Liang Lu, Xingxing Zhang, and Steve Renais, · 2016
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei et al., · 2016
Cited alongside, same era.
“On online attention-based speech recognition and joint Mandarin character-Pinyin training,”
William Chan and Ian Lane, · 2016
Cited alongside, same era.
“Conditional image generation with PixelCNN decoders,”
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al., · 2016
Cited alongside, same era.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
“Automatic differentiation in PyTorch,”
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, · 2017
Later among the works it cites.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Later among the works it cites.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Later among the works it cites.
“Self-attentional acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
“An analysis of ”attention” in sequence-to-sequence models,”
Rohit Prabhavalkar, Tara N Sainath, Bo Li, Kanishka Rao, and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Acoustic-to-word model without OOV,”
Jinyu Li, Guoli Ye, Rui Zhao, Jasha Droppo, and Yifan Gong, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2017
Cited alongside, same era.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, ZhiJeng Chen, and Rohit Prabhavalkar, · 2018
Later among the works it cites.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Closest in time.
“The Speechtransformer for large-scale Mandarin Chinese speech recognition,”
Yuanyuan Zhao, Jie Li, Xiaorui Wang, and Yan Li, · 2019
Closest in time.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Closest in time.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Closest in time.
“Generating long sequences with sparse transformers,”
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever, · 2019
Closest in time.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2019
Closest in time.