Fetching the paper…
Reading the bibliography…
The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency.
“Response time in man-computer conversational transactions,”
R. B. Miller, · 1968
Earlier work this paper cites.
The Harpy Speech Recognition System
B. T. Lowerre, · 1976
Earlier work this paper cites.
“A comparison of several approximate algorithms for finding multiple (N-best) sentence hypotheses,”
R. Schwartz and S. Austin, · 1991
Earlier work this paper cites.
“A word graph algorithm for large vocabulary continuous speech recognition,”
S. Ortmanns, H. Ney, and X. Aubert, · 1997
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep neural networks,”
A. Graves, A.-R. Mohamed, and G. Hinton, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Earlier work this paper cites.
“Bringing contextual information to Google speech recognition,”
P. Aleksic, M. Ghodsi, A. Michaely, C. Allauzen, K. Hall, B. Roark, D. Rybach, and P. Moreno, · 2015
Earlier work this paper cites.
“From Feedforward to Recurrent LSTM Neural Networks for Language Models,”
M. Sundermeyer, H. Ney, and R. Schlüter, · 2015
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” Available online: http://download.tensorflow.org/paper/whitepaper2015.pdf, 2015
M. Abadi et al., · 2015
Cited alongside, same era.
“Lower frame rate neural network acoustic models,”
G. Pundak and T. N. Sainath, · 2016
Cited alongside, same era.
“Personalized speech recognition on mobile devices,”
I. McGraw, R. Prabhabalkar, R. Alvarez, M. Gonzalez, K. Rao, D. Rybach, O. Alsharif, H. Sak, A. Gruenstein, F. Beaufays, and C. Parada, · 2016
Cited alongside, same era.
“Two Efficient Lattice Rescoring Methods Using Recurrent Neural Network Language Models,”
X. Liu, X. Chen, Y. Wang, M. J. F. Gales, and P. C. Woodland, · 2016
Cited alongside, same era.
“An analysis of ”attention” in sequence-to-sequence models,”
R. Prabhavalkar, T. N. Sainath, B. Li, K. Rao, and N. Jaitly, · 2017
Later among the works it cites.
“Towards Better Decoding and Language Model Integration in Sequence to Sequence Models,”
J. K. Chorowski and N. Jaitly, · 2017
Later among the works it cites.
“Generated of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, N. Jaitly, B. Li, and J. Chorowski, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Gonzalvo, S. Tazari, C. Chan, M. Becker, A. Gutkin, and H. Silen, · 2016
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Cited alongside, same era.
“Monotonic chunkwise alignments,”
C.-C. Chiu and C. Raffel, · 2017
Cited alongside, same era.
“Lattice Rescoring Strategies for Long Short Term Memory Language Models in Speech Recognition,”
S. Kumar, M. Nirschl, D. Holtmann-Rice, H. Liao, A. T. Suresh, and F. Yu, · 2017
Cited alongside, same era.
Later among the works it cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, · 2018
Later among the works it cites.
“Minimum Word Error Rate Training for Attention-based Sequence-to-sequence Models,”
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C. C. Chiu, and A. Kannan, · 2018
Later among the works it cites.
“Deep Context: End-to-End Contextual Speech Recognition,”
G. Pundak, T. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, · 2018
Later among the works it cites.
“Streaming End-to-end Speech Recognition For Mobile Devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K. Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein, · 2019
Closest in time.
“Lingvo: a modular and scalable framework for sequence-to-sequence modeling,” 2019
Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, et al., · 2019
Closest in time.