Fetching the paper…
Reading the bibliography…
The attention mechanism of the Listen, Attend and Spell (LAS) model requires the whole input sequence to calculate the attention context and thus is not suitable for online speech recognition.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Lower frame rate neural network acoustic models,”
Golan Pundak and Tara N Sainath, · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, and et al., · 2016
Earlier work this paper cites.
“Online and linear-time attention by enforcing monotonic alignments,”
Colin Raffel, Douglas Eck, Peter J Liu, and et al., · 2017
Earlier work this paper cites.
“Sequence-to-sequence models can directly translate foreign speech,”
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, and et al., · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Bu Hui, Du Jiayu, Na Xingyu, and et al., · 2017
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chungcheng Chiu, Tara N Sainath, Yonghui Wu, and et al., · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chungcheng Chiu and Colin Raffel, · 2018
Cited alongside, same era.
“Extending recurrent neural aligner for streaming end-to-end speech recognition in mandarin,”
Linhao Dong, Shiyu Zhou, Wei Chen, and Bo Xu, · 2018
Cited alongside, same era.
“Minimum word error rate training for attention-based sequence-to-sequence models,”
Rohit Prabhavalkar, Tara N Sainath, Yonghui Wu, and et al., · 2018
Cited alongside, same era.
“Extending recurrent neural aligner for streaming end-to-end speech recognition in mandarin,”
Linhao Dong, Shiyu Zhou, Wei Chen, and Bo Xu, · 2018
Cited alongside, same era.
“Output-gate projected gated recurrent unit for speech recognition.,”
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, and et al., · 2019
Later among the works it cites.
“Model unit exploration for sequence-to-sequence speech recognition,”
Kazuki Irie, Rohit Prabhavalkar, Anjuli Kannan, and et al., · 2019
Later among the works it cites.
“Lingvo: a modular and scalable framework for sequence-to-sequence modeling,”
Jonathan Shen, Patrick Nguyen, Yonghui Wu, and et al., · 2019
Later among the works it cites.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, and et al., · 2019
Later among the works it cites.
“Adversarial regularization for attention based end-to-end robust speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gaofeng Cheng, Daniel Povey, Lu Huang, and et al., · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, and et al., · 2019
Cited alongside, same era.
Sun Sining, Guo Pengcheng, Xie Lei, and Hwang Mei-Yuh, · 2019
Later among the works it cites.
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengkun Tian, and Zhengqi Wen, · 2019
Later among the works it cites.
“The speechtransformer for large-scale mandarin chinese speech recognition,”
Jie Li, Xiaorui Wang, Yan Li, et al., · 2019
Later among the works it cites.