Fetching the paper…
Reading the bibliography…
The Listen, Attend and Spell (LAS) model and other attention-based automatic speech recognition (ASR) models have known limitations when operated in a fully online mode.
“Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,”
S. Kanthak and H. Ney, · 2002
Earlier work this paper cites.
“Grapheme based speech recognition,”
Mirjam Killer, Sebastian Stüker, and Tanja Schultz, · 2003
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukás Burget, O. Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlíček, Yanmin Qian, Petr Schwarz, Jan Silovský, Georg Stemmer, and Karel Veselý, · 2011
Earlier work this paper cites.
“Deep Neural Networks for Acoustic Modeling in Speech Recognition,”
G. Hinton, L. Deng, D. Yu, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, G. Dahl, and B. Kingsbury, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Developing speech recognition systems for corpus indexing under the IARPA Babel program,”
Jia Cui, Xiaodong Cui, Bhuvana Ramabhadran, Janice Kim, Brian Kingsbury, Jonathan Mamou, Lidia Mangu1, Michael Picheny, Tara N. Sainath, and Abhinav Sethy, · 2013
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering,”
Kai Chen and Qiang Huo, · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna, · 2016
Cited alongside, same era.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“Hybrid CTC/Attention Architecture for End-to-End Speech Recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Online and linear-time attention by enforcing monotonic alignments,”
Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, , and Douglas Eck, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Improving RNN Transducer Modeling for End-to-End Speech Recognition,”
Jinyu Li, Rui Zhao, Hu Hu, and Yifan Gong, · 2019
Later among the works it cites.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2019
Later among the works it cites.
“An analysis of local monotonic attention variants,”
André Merboldt, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Later among the works it cites.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N. Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, Chung-Cheng Chiu, David Garcia, Alex Gruenstein, Ke Hu, Minho Jin, Anjuli Kannan, Qiao Liang, Ian McGraw, Cal Peyser, Rohit Prabhavalkar, Golan Pundak, David Rybach, Yuan Shangguan, Yash Sheth, Trevor Strohman, Mirkó Visontai, Yonghui Wu, Yu Zhang, and Ding Zhao, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chung-Cheng Chiu and Colin Raffel, · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo yiin Chang, Kanishka Rao, and Alexander Gruenstein, · 2019
Cited alongside, same era.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Cited alongside, same era.
Connectionist Speech Recognition: A Hybrid Approach
Hervé Bourlard and Nelson Morgan,
Cited in the paper.
Closest in time.
“An attention-based joint acoustic and text on-device end-to-end model,”
Tara N. Sainath, Ruoming Pang, Ron J. Weiss, Yanzhang He, Chung cheng Chiu, and Trevor Strohman, · 2020
Closest in time.
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo yiin Chang, Tara N. Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Closest in time.
“Minimum latency training strategies for streaming sequence-to-sequence asr,”
Hirofumi Inaguma, Yashesh Gaur, Liang Lu, Jinyu Li, and Yifan Gong, · 2020
Closest in time.