Fetching the paper…
Reading the bibliography…
Recently, online end-to-end ASR has gained increasing attention.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Feedforward sequential memory networks: A new structure to learn long-term dependency,”
Shiliang Zhang, Cong Liu, Hui Jiang, Si Wei, Lirong Dai, and Yu Hu, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Online and linear-time attention by enforcing monotonic alignments,”
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck, · 2017
Earlier work this paper cites.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2017
Cited alongside, same era.
“Improving the performance of online neural transducer models,”
Tara N Sainath, Chung-Cheng Chiu, Rohit Prabhavalkar, Anjuli Kannan, Yonghui Wu, Patrick Nguyen, and ZhiJeng Chen, · 2018
Cited alongside, same era.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2018
Cited alongside, same era.
“Deep-FSMN for large vocabulary continuous speech recognition,”
Shiliang Zhang, Lei Ming, Zhijie Yan, and Lirong Dai, · 2018
Cited alongside, same era.
“Aishell-2: transforming mandarin asr research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Cited alongside, same era.
“Investigation of modeling units for mandarin speech recognition using DFSMN-CTC-sMBR,”
Shiliang Zhang, Ming Lei, Yuan Liu, and Wei Li, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.
“Streaming automatic speech recognition with the transformer model,”
Niko Moritz, Takaaki Hori, and Jonathan Le, · 2020
Closest in time.
“Streaming chunk-aware multihead attention for online end-to-end speech recognition,”
Shiliang Zhang, Zhifu Gao, Haoneng Luo, Ming Lei, Jie Gao, Zhijie Yan, and Lei Xie, · 2020
Closest in time.
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo-yiin Chang, Tara N Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2018
Cited alongside, same era.
“Online hybrid ctc/attention architecture for end-to-end speech recognition,”
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, and Yonghong Yan, · 2019
Cited alongside, same era.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Cited alongside, same era.
“Two-pass end-to-end speech recognition,”
Tara N Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, et al., · 2019
Cited alongside, same era.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, et al., · 2020
Closest in time.
“San-m: Memory equipped self-attention for end-to-end speech recognition,”
Zhifu Gao, Shiliang Zhang, Ming Lei, and Ian McLoughlin, · 2020
Closest in time.
“Cif: Continuous integrate-and-fire for end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Closest in time.