Fetching the paper…
Reading the bibliography…
We develop streaming keyword spotting systems using a recurrent neural network transducer (RNN-T) model: an all-neural, end-to-end trained, sequence-to-sequence model which jointly learns acoustic and language model components.
“A hidden markov model based keyword recognition system,”
R. C. Rose and D. B. Paul, · 1990
Earlier work this paper cites.
“Keyword-spotting using sri’s decipher large-vocabulary speech-recognition system,”
Mitchel Weintraub, · 1993
Earlier work this paper cites.
““long short-term memory”,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Confidence intervals for the area under the roc curve,”
C. Cortes and M. Mohri, · 2004
Earlier work this paper cites.
Phoneme Based Acoustics Keyword Spotting in Informal Continuous Speech
I. Szöke, P. Schwarz, P. Matějka, L. Burget, M. Karafiát, and J. Černocký, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: Labeling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Results of the 2006 spoken term detection evaluation,”
J. G. Fiscus, J. Ajot, J. S. Garofolo, and G. Doddingtion, · 2007
Earlier work this paper cites.
“Rapid and accurate spoken term detection,”
D. R. H. Miller, M. Kleber, C.-L. Kao, O. Kimball, T. Colthurst, S. A. Lowe, R. M. Schwartz, and H. Gish, · 2007
Earlier work this paper cites.
“The SRI/OGI 2006 spoken term detection system,”
D. Vergyri, I. Shafran, A. Stolcke, R. R. Gadde, M. Akbacak, B. Roark, and W. Wang, · 2007
Earlier work this paper cites.
“An application of recurrent neural networks to discriminative keyword spotting,”
S. Fernández, A. Graves, and J. Schmidhuber, · 2007
Earlier work this paper cites.
“Discriminative keyword spotting,”
J. Keshet, D. Grangier, and S. Bengio, · 2009
Earlier work this paper cites.
“Query-by-example spoken term detection using phonetic posteriorgram templates,”
T. J. Hazen, W. Shen, and C. M. White, · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“White listing and score normalization for keyword spotting of noisy speech,”
B. Zhang, R. Schwartz, S. Tsakalidis, L. Nguyen, and S. Matsoukas, · 2012
Earlier work this paper cites.
“Large Scale Distributed Deep Networks,”
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng, · 2012
Earlier work this paper cites.
“Discriminative articulatory models for spoken term detection in low-resource conversational settings,”
R. Prabhavalkar, K. Livescu, E. Fosler-Lussier, and J. Keshet, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A.-R. Mohamed, and G. E. Hinton, · 2013
Cited alongside, same era.
“Context dependent acoustic keyword spotting using deep neural network,”
G. Wang and K. C. Sim, · 2013
Cited alongside, same era.
“Small footprint keyword spotting using deep neural networks,”
G. Chen, C. Parada, and G. Heigold, · 2014
Cited alongside, same era.
“Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks,”
R. Prabhavalkar, R. Alvarez, C. Parada, P. Nakkiran, and T. N. Sainath, · 2015
Cited alongside, same era.
“Convolutional neural networks for small footprint keyword spotting,”
T. N. Sainath and C. Parada, · 2015
Cited alongside, same era.
“Query-by-example keyword spotting using long short-term memory networks,”
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Later among the works it cites.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Later among the works it cites.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
L. Lu, X. Zhang, and S. Renals, · 2016
Later among the works it cites.
“Unrestricted vocabulary keyword spotting using lstm-ctc,”
Yimeng Zhuang, Xuankai Chang, Yanmin Qian, and Kai Yu, · 2016
Later among the works it cites.
“Personalized speech recognition on mobile devices,”
I. McGraw, R. Prabhavalkar, R. Alvarez, M. Gonzalez Arenas, K. Rao, D. Rybach, O. Alsharif, H. Sak, A. Gruenstein, F. Beaufays, and C. Parada, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Chen, C. Parada, and T. N. Sainath, · 2015
Cited alongside, same era.
“Online keyword spotting with a character-level recurrent neural network,”
K. Hwang, M. Lee, and W. Sung, · 2015
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
H. Sak, A. W. Senior, K. Rao, and F. Beaufays, · 2015
Cited alongside, same era.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Cited alongside, same era.
“A keyword-aware grammar framework for lvcsr-based spoken keyword search,”
I.-F. Chen, C. Ni, B. P. Lim, N. F. Chen, and C.-H. Lee, · 2015
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” Available online: http://download.tensorflow.org/paper/whitepaper2015.pdf, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M.Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mane, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viegas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, · 2015
Cited alongside, same era.
“Model compression applied to small-footprint keyword spotting,”
G. Tucker, M. Wu, M. Sun, S. Panchapagesan, G. Fu, and S. Vitaladevuni, · 2016
Cited alongside, same era.
R. Alvarez, R. Prabhavalkar, and A. Bakhtin, · 2016
Later among the works it cites.
“On the compression of recurrent neural networks with an application to lvcsr acoustic modeling for embedded speech recognition,”
R. Prabhavalkar, O. Alsharif, A. Bruguier, and L. McGraw, · 2016
Later among the works it cites.
“Lstm: A search space odyssey,”
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber, · 2016
Later among the works it cites.
“Convolutional recurrent neural networks for small-footprint keyword spotting,”
S. Ö. Arık, M. Kliegl, R. Child, J. Hestness, A. Gibiansky, C. Fougner, R. Prenger, and A. Coates, · 2017
Closest in time.
“Acoustic modeling for google home,”
B. Li, T. N. Sainath, J. Caroselli, A. Narayanan, M. Bacchiani, A. Misra, I. Shafran, H. Sak, G. Pundak, K. Chin, K. Sim, R. J. Weiss, K. W. Wilson, E. Variani, C. Kim, O. Siohan, M. Weintraub, E. McDermott, R. Rose, and M. Shannon, · 2017
Closest in time.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence-to-sequence mapping,”
H. Sak, M. Shannon, K. Rao, and F. Beaufays, · 2017
Closest in time.
“A comparison of sequence-to-sequence models for speech recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Closest in time.
“An analysis of “attention” in sequence-to-sequence models,”
R. Prabhavalkar, T. N. Sainath, B. Li, K. Rao, and N. Jaitly, · 2017
Closest in time.
“End-to-end speech recognition and keyword search on low-resource languages,”
A. Rosenberg, K. Audhkhasi, A. Sethy, B. Ramabhadran, and M. Picheny, · 2017
Closest in time.
“End-to-end asr-free keyword search from speech,”
K. Audhkhasi, A. Rosenberg, A. Sethy, B. Ramabhadran, and B. Kingsbury, · 2017
Closest in time.
“Optimizing expected word error rate via sampling for speech recognition,”
M. Shannon, · 2017
Closest in time.