Fetching the paper…
Reading the bibliography…
The Transformer self-attention network has shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR) systems.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K. Paliwal, · 1997
Earlier work this paper cites.
“Spontaneous speech corpus of Japanese,”
Kikuo Maekawa, Hanae Koiso, Sadaoki Furui, and Hitoshi Isahara, · 2000
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“HKUST/MTS: A very large scale Mandarin telephone speech corpus,”
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff, · 2006
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-Rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Learning small-size DNN with output-distribution-based criteria,”
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong, · 2014
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K. Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Deep Speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei et al., · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“An online sequence-to-sequence model using partial conditioning,”
Navdeep Jaitly, Quoc V Le, Oriol Vinyals, Ilya Sutskever, David Sussillo, and Samy Bengio, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2017
Cited alongside, same era.
“Knowledge distillation for small-footprint highway networks,”
Liang Lu, Michelle Guo, and Steve Renals, · 2017
Cited alongside, same era.
“AIShell-1: An open-source Mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Towards online end-to-end transformer automatic speech recognition,”
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura, and Shinji Watanabe, · 2019
Later among the works it cites.
“Generating long sequences with sparse transformers,”
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever, · 2019
Later among the works it cites.
“Vectorized beam search for ctc-attention-based speech recognition,”
Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Niko Moritz, and Jonathan Le Roux, · 2019
Later among the works it cites.
“Self-attention transducers for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, and Zhengqi Wen, · 2019
Later among the works it cites.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel, · 2018
Cited alongside, same era.
“A time-restricted self-attention layer for ASR,”
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Cited alongside, same era.
“The SpeechTransformer for large-scale Mandarin Chinese speech recognition,”
Yuanyuan Zhao, Jie Li, Xiaorui Wang, and Yan Li, · 2019
Cited alongside, same era.
“A comparative study on transformer vs RNN in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
Cited alongside, same era.
“Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
Cited alongside, same era.
“Transformer-XL: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Later among the works it cites.
“Streaming automatic speech recognition with the transformer model,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2020
Closest in time.
“Transformer online CTC/attention end-to-end speech recognition architecture,”
Haoran Miao, Gaofeng Cheng, Zhang Pengyuan, and Yonghong Yan, · 2020
Closest in time.
“Minimum latency training strategies for streaming sequence-to-sequence ASR,”
Hirofumi Inaguma, Yashesh Gaur, Liang Lu, Jinyu Li, and Yifan Gong, · 2020
Closest in time.
“Synchronous transformers for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Ye Bai, Jianhua Tao, Shuai Zhang, and Zhengqi Wen, · 2020
Closest in time.
“CIF: Continuous integrate-and-fire fore end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Closest in time.
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu, · 2020
Closest in time.
“Alignment-length synchronous decoding for RNN transducer,”
George Saon, Zoltán Tüske, and Kartik Audhkhasi, · 2020
Closest in time.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Closest in time.