Fetching the paper…
Reading the bibliography…
The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR) systems.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K. Paliwal, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-Rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K. Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
Navdeep Jaitly, David Sussillo, Quoc V. Le, Oriol Vinyals, Ilya Sutskever, and Samy Bengio, · 2015
Earlier work this paper cites.
“Deep Speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei et al., · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
Liang Lu, Xingxing Zhang, and Steve Renais, · 2016
Earlier work this paper cites.
“On online attention-based speech recognition and joint Mandarin character-Pinyin training,”
William Chan and Ian Lane, · 2016
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, ZhiJeng Chen, and Rohit Prabhavalkar, · 2018
Later among the works it cites.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Closest in time.
“The Speechtransformer for large-scale Mandarin Chinese speech recognition,”
Yuanyuan Zhao, Jie Li, Xiaorui Wang, and Yan Li, · 2019
Closest in time.
“A comparative study on transformer vs RNN in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chung-Cheng Chiu and Colin Raffel, · 2017
Cited alongside, same era.
“AIShell-1: An open-source Mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Cited alongside, same era.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Self-attentional acoustic models,”
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Cited alongside, same era.
Closest in time.
“Transformer ASR with contextual block processing,”
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura, and Shinji Watanabe, · 2019
Closest in time.
“An analysis of local monotonic attention variants,”
André Merboldt, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“Online hybrid CTC/attention architecture for end-to-end speech recognition,”
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, and Yonghong Yan, · 2019
Closest in time.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2019
Closest in time.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Closest in time.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2019
Closest in time.