Fetching the paper…
Reading the bibliography…
Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic speech recognition (ASR).
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Povey Daniel, Ghoshal Arnab, Boulianne Gilles, Burget Lukas, Glembek Ondrej, Goel Nagendra, Hannemann Mirko, Motlicek Petr, Qian Yanmin, Schwarz Petr, Silovsky Jan, Stemmer Georg, and Vesely Karel, · 2011
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Earlier work this paper cites.
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Earlier work this paper cites.
“AISHELL-2: transforming mandarin ASR research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Earlier work this paper cites.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Xiaofei Wang, and Shinji Watanabe, · 2019
Cited alongside, same era.
“Listen and fill in the missing letters: Non-autoregressive transformer for speech recognition,”
Nanxin Chen, Shinji Watanabe, Jesús Villalba, and Najim Dehak, · 2019
Cited alongside, same era.
“Mask-predict: Parallel decoding of conditional masked language models,”
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer, · 2019
Cited alongside, same era.
“Hard but robust, easy but sensitive: How encoder and decoder perform in neural machine translation,”
Tianyu He, Xu Tan, and Tao Qin, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“Ectc-docd: An end-to-end structure with CTC encoder and OCD decoder for speech recognition,”
Cheng Yi, Feng Wang, and Bo Xu, · 2019
Later among the works it cites.
“Listen attentively, and spell once: Whole sentence generation via a non-autoregressive architecture for low-latency speech recognition,”
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengkun Tian, Zhengqi Wen, and Shuai Zhang, · 2020
Closest in time.
“Spike-triggered non-autoregressive transformer for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, Shuai Zhang, and Zhengqi Wen, · 2020
Closest in time.
“Mask CTC: non-autoregressive end-to-end ASR with CTC and mask predict,”
Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, and Tetsunori Kobayashi, · 2020
Closest in time.
“Semi-autoregressive training improves mask-predict decoding,”
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su, Min Luo, Dong Yu, and Lei Xie, · 2019
Cited alongside, same era.
“Adversarial regularization for attention based end-to-end robust speech recognition,”
Sining Sun, Pengcheng Guo, Lei Xie, and Mei-Yuh Hwang, · 2019
Cited alongside, same era.
“A study of non-autoregressive model for sequence generation,”
Yi Ren, Jinglin Liu, Xu Tan, Zhou Zhao, Sheng Zhao, and Tie-Yan Liu, · 2020
Closest in time.
“CIF: continuous integrate-and-fire for end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Closest in time.