Fetching the paper…
Reading the bibliography…
Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and a substantial part of NAR transformer models is to use the casual mask to limit token dependencies.
“A tutorial on the cross-entropy method,”
Pieter-Tjerk de Boer, Dirk P.Kroese, Shie Mannor, and Reuven Y.Rubinstein, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton, · 2016
Earlier work this paper cites.
“Rethinking the Inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, et al., · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and aspeech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Speech-Transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, et al., · 2018
Earlier work this paper cites.
“Non-autoregressive neural machine translation,”
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher, · 2018
Cited alongside, same era.
“Deterministic non-autoregressive neural sequence modeling by iterative refinement,”
Jason Lee, Elman Mansimov, and Kyunghyun Cho, · 2018
Cited alongside, same era.
“Espnet: end-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, et al., · 2018
Cited alongside, same era.
“Scaling neural machine translation,”
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli, · 2018
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, et al., · 2019
Cited alongside, same era.
“Speech transformer with speaker aware persistent memory,”
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Joty Shafiq, Eng Siong Chng, and Bin Ma, · 2020
“Listen attentively, and spell once: Whole sentence generation via a non-autoregressive architecture for low-latency speech recognition,”
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengkun Tian, Zhengqi Wen, et al., · 2020
Later among the works it cites.
“Transformer with bidirectional decoder for speech recognition,”
Xi Chen, Songyang Zhang, Dandan Song, Peng Ouyang, and Shouyi Yin, · 2020
Later among the works it cites.
“Non-autoregressive machine translation with disentangled contexts transformer,”
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu, · 2020
Later among the works it cites.
“Insertion-based modeling for end-to-end automatic speech recognition,”
Yuya Fujita, Shinji Watanabe, Motoi Omachi, and Xuankai Chang, · 2020
Later among the works it cites.
“Listen and fill in the missing letters: Non-autoregressive transformer for speech recognition,”
Nanxin Chen, Shinji Watanabe, Jesús Villalba, and Najim Dehak, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo-Yiin Chang, Tara N. Sainath, Ruoming Pang, Yanzhang He, et al., · 2020
Cited alongside, same era.
“Aligned cross entropy for non-autoregressive machine translation,”
Marjan Ghazvininejad, VladimirLuke Karpukhin, Luke Zettlemoyer, and Omer Levy, · 2020
Cited alongside, same era.
“Spike-triggered non-autoregressive transformer for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, Shuai Zhang, et al., · 2020
Cited alongside, same era.
“Mask CTC: non-autoregressive end-to-end asr with CTC and Mask predict,”
Yosuke Higuchi, Shinji Watanabe, Chen Nanxin, Tetsuji Ogawa, et al., · 2020
Cited alongside, same era.
“Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration,”
Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe, Marc Delcroix, Atsunori Ogawa, and Tomohiro Nakatani, · 2020
Later among the works it cites.
“CASS-NAT: CTC alignment-based single step non-autoregressive transformer for speech recognition,”
Ruchao Fan, Wei Chu, Peng Chang, and Jing Xiao, · 2021
Closest in time.
“Non-autoregressive transformer asr with CTC-Enhanced decoder input,”
Xingchen Song, Zhiyong Wu, Yiheng Huang, Chao Weng, Dan Su, et al., · 2021
Closest in time.
“TSNAT: two-step non-autoregressvie transformer models for speech recognition,”
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, Shuai Zhang, Zhengqi Wen, and Xuefei Liu, · 2021
Closest in time.
“U2++: Unified two-pass bidirectional end-to-end model for speech recognition,”
Di Wu, Binbin Zhang, Chao Yang, Zhendong Peng, Wenjing Xia, et al., · 2021
Closest in time.