Fetching the paper…
Reading the bibliography…
End-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance compared to conventional systems.
“Connectionist Temporal Classification: Labeling Unsegmented Seuqnece Data with Recurrent Neural Networks,”
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription,”
Hank Liao, Erik McDermott, and Andrew Senior, · 2013
Earlier work this paper cites.
“Generating sequences with recurrent neural networks,” 2013
Alex Graves, · 2013
Earlier work this paper cites.
“Towards End-to-End Speech Recognition with Recurrent Neural Networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“End-to-End Attention-based Large Vocabulary Speech Recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Cited alongside, same era.
“Exploring Neural Transducers for End-to-End Speech Recognition,”
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur, Yi Li, Hairong Liu, Sanjeev Satheesh, Anuroop Sriram, and Zhenyao Zhu, · 2017
Cited alongside, same era.
“A Comparison of Sequence-to-sequence Models for Speech Recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Cited alongside, same era.
“Very Deep Convolutional Networks for End-to-End Speech Recognition,”
Y. Zhang, W. Chan, and N. Jaitly, · 2017
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
Hagen Soltau, Hank Liao, and Hasim Sak, · 2017
Cited alongside, same era.
“Online and linear-time attention by enforcing monotonic alignments,”
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
Later among the works it cites.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2018
Later among the works it cites.
“Toward domain-invariant speech recognition via large scale training,”
Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed Elfeky, Parisa Haghani, Trevor Strohman, and Michiel Bacchiani, · 2018
Later among the works it cites.
“Compression of end-to-end models,”
Ruoming Pang, Tara Sainath, Rohit Prabhavalkar, Suyog Gupta, Yonghui Wu, Shuyuan Zhang, and Chung-Cheng Chiu, · 2018
Later among the works it cites.
“Efficient implementation of recurrent neural network transducer in tensorflow,”
Tom Bagby, Kanishka Rao, and Khe Chai Sim, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, and Douglas Eck, · 2017
Cited alongside, same era.
“Local monotonic attention mechanism for end-to-end speech and language processing,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Cited alongside, same era.
“Improving the efficiency of forward-backward algorithm using batched computation in tensorflow,”
Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, and Michiel Bacchiani, · 2017
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo yiin Chang, Kanishka Rao, and Alexander Gruenstein, · 2019
Closest in time.
“Monotonic infinite lookback attention for simultaneous machine translation,”
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, and Colin Raffel, · 2019
Closest in time.
“Lingvo: a modular and scalable framework for sequence-to-sequence modeling,” 2019
Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, and et al., · 2019
Closest in time.