Fetching the paper…
Reading the bibliography…
End-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition.
“Long Short-Term Memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Weighted finite-state transducers in speech recognition,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2002
Earlier work this paper cites.
“Speechalator: Two-way speech-to-speech translation on a consumer PDA,”
A. Waibel et al., · 2003
Earlier work this paper cites.
“Connectionist Temporal Classification: Labeling Unsegmented Sequenece Data with Recurrent Neural Networks,”
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Embedded speech recognition applications in mobile phones: Status, trends, and challenges,”
J. Cohen, · 2008
Earlier work this paper cites.
“Your Word is my Command”: Google Search by Voice: A Case Study
J. Schalkwyk, D. Beeferman, F. Beaufays, B. Byrne, C. Chelba, M. Cohen, M. Kamvar, and B. Strope, · 2010
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep neural networks,”
A. Graves, A.-R. Mohamed, and G. Hinton, · 2012
Earlier work this paper cites.
“Japanese and korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
Mike Schuster and Kaisuke Nakajima, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling,”
H. Sak, A. Senior, and F. Beaufays, · 2014
Earlier work this paper cites.
“Convolutional neural networks for small-footprint keyword spotting,”
T. N. Sainath and C. Parada, · 2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Earlier work this paper cites.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Cited alongside, same era.
“Bringing Contextual Information to Google Speech Recognition,”
P. Aleksic, M. Ghodsi, A. Michaely, C. Allauzen, K. Hall, B. Roark, D. Rybach, and P. Moreno, · 2015
Cited alongside, same era.
“Composition-based on-the-fly rescoring for salient n-gram biasing,”
K.B. Hall, E. Cho, C. Allauzen, F. Beaufays, N. Coccaro, K. Nakajima, M. Riley, B. Roark, D. Rybach, and L. Zhang, · 2015
Cited alongside, same era.
“Sequence-based class tagging for robust transcription in asr,”
L. Vasserman, V. Schogol, and K.B. Hall, · 2015
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” Available online: http://download.tensorflow.org/paper/whitepaper2015.pdf, 2015
M. Abadi et al., · 2015
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
H. Soltau, H. Liao, and H. Sak, · 2017
Later among the works it cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Later among the works it cites.
“Monotonic chunkwise alignments,”
C.-C. Chiu and C. Raffel, · 2017
Later among the works it cites.
“Improving the Efficiency of Forward-Backward Algorithm using Batched Computation in TensorFlow,”
K. Sim, A. Narayanan, T. Bagby, T.N. Sainath, and M. Bacchiani, · 2017
Later among the works it cites.
“Parallel wavenet: Fast high-fidelity speech synthesis,”
A. van den Oord, Y. Li, and I. Babuschkin et. al., · 2017
Later among the works it cites.
“Reducing the Computational Complexity for Whole Word Models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Personalized speech recognition on mobile devices,”
I. McGraw, R. Prabhavalkar, R. Alvarez, M. G. Arenas, K. Rao, D. Rybach, O. Alsharif, H. Sak, A. Gruenstein, F. Beaufays, and C. Parada, · 2016
Cited alongside, same era.
“End-to-End Attention-based Large Vocabulary Speech Recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Cited alongside, same era.
J. L. Ba, R. Kiros, and G. E. Hinton, · 2016
Cited alongside, same era.
“On the efficient representation and execution of deep acoustic models,”
R. Alvarez, R. Prabhavalkar, and A. Bakhtin, · 2016
Cited alongside, same era.
“Lower Frame Rate Neural Network Acoustic Models,”
G. Pundak and T. N. Sainath, · 2016
Cited alongside, same era.
“Recent Advances in Google Real-time HMM-driven Unit Selection Synthesizer,”
X. Gonzalvo, S. Tazari, C. Chan, M. Becker, A. Gutkin, and H. Silen, · 2016
Cited alongside, same era.
“Convolutional recurrent neural networks for small-footprint keyword spotting,”
S. O. Arik, M. Kliegl, R. Child, J. Hestness, A. Gibiansky, C. Fougner, R. Prenger, and A. Coates, · 2017
Cited alongside, same era.
H. Soltau, H. Liao, and H. Sak, · 2017
Later among the works it cites.
“In-datacenter performance analysis of a tensor processing unit,”
N. P. Jouppi et al., · 2017
Later among the works it cites.
“An RNN Model of Text Normalization,”
R. Sproat and N. Jaitly, · 2017
Later among the works it cites.
“Generated of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C. C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, N. Jaitly, B. Li, and J. Chorowski, · 2018
Closest in time.
“Contextual speech recognition in end-to-end neural network systems using beam search,”
I. Williams, A. Kannan, P. Aleksic, D. Rybach, and T. N. Sainath, · 2018
Closest in time.
“Deep Context: End-to-End Contextual Speech Recognition,”
G. Pundak, T. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, · 2018
Closest in time.
“Introducing the Model Optimization Toolkit for TensorFlow,” https://medium.com/tensorflow/introducing-the-model-optimization-toolkit-for-tensorflow-254aca1ba0a3,
R. Alvarez, R. Krishnamoorthi, S. Sivakumar, Y. Li, A. Chiao, P Warden, S. Shekhar, S. Sirajuddin, and Davis. T., · 2018
Closest in time.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, · 2018
Closest in time.