Fetching the paper…
Reading the bibliography…
Accurate on-device keyword spotting (KWS) with low false accept and false reject rate is crucial to customer experience for far-field voice control of conversational agents.
J. G. Wilpon, L. R. Rabiner, C.-H. Lee, and E. R. Goldman, “Automatic recognition of keywords in unconstrained speech using hidden markov models,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 38, no. 11, pp. 1870–1878, 1990
1990
Earlier work this paper cites.
R. C. Rose and D. B. Paul, “A hidden markov model based keyword recognition system,” in International Conference on Acoustics, Speech, and Signal Processing , 1990, pp. 129–132
1990
Earlier work this paper cites.
J. Wilpon, L. Miller, and P. Modi, “Improvements and applications for key word recognition using hidden markov modeling techniques,” in [Proceedings] ICASSP 91: 1991 International Conference on Acoustics, Speech, and Signal Processing , 1991, pp. 309–312
1991
Earlier work this paper cites.
M. Weintraub, “Improved keyword-spotting using sri’s decipher™ large-vocabuarly speech-recognition system,” in HLT ’93 Proceedings of the workshop on Human Language Technology , 1993, pp. 114–118
1993
Earlier work this paper cites.
M. Gales, “Maximum likelihood linear transformations for hmm-based speech recognition☆,” Computer Speech & Language , vol. 12, no. 2, pp. 75–98, 1998
1998
Earlier work this paper cites.
E. Hänsler and G. Schmidt, Acoustic Echo and Noise Control: A Practical Approach , 2004
2004
Earlier work this paper cites.
O. Kalinli, N. L. Seltzer, and A. Acero, “Noise adaptive training using a vector taylor series approach for noise robust automatic speech recognition,” in Acoustics, Speech and Signal Processing, 2009. ICASSP 2009. IEEE International Conference on . IEEE, 2009, pp. 3825–3828
2009
Earlier work this paper cites.
K. Kumatani, J. W. McDonough, and B. Raj, “Microphone array processing for distant speech recognition: From close-talking microphones to far-field sensors,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 127–140, 2012
2012
Earlier work this paper cites.
T. Virtanen, R. Singh, and B. Raj, “Techniques for noise robustness in automatic speech recognition,” 2012
2012
Cited alongside, same era.
M. L. Seltzer, “Acoustic model training for robust speech recognition,” Techniques for Noise Robustness in Automatic Speech Recognition , pp. 347–368, 2012
2012
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. rahman Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Cited alongside, same era.
C. Paleologu, J. Benesty, and S. Ciochina, “Study of the general kalman filter for echo cancellation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 21, no. 8, pp. 1539–1549, 2013
2013
Cited alongside, same era.
R. Hsiao, J. Ma, W. Hartmann, M. Karafiat, F. Grezl, L. Burget, I. Szoke, J. H. Cernocky, S. Watanabe, Z. Chen, S. H. Mallidi, H. Hermansky, S. Tsakalidis, and R. Schwartz, “Robust speech recognition in unknown reverberant and noisy conditions,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015, pp. 533–538
2015
Later among the works it cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proceedings of INTERSPEECH , 2015
2015
Later among the works it cites.
N. Strom, “Scalable distributed dnn training using commodity gpu cloud computing.” in INTERSPEECH , 2015, pp. 1488–1492
2015
Later among the works it cites.
M. Sun, A. Raju, G. Tucker, S. Panchapagesan, G. Fu, A. Mandal, S. Matsoukas, N. Strom, and S. Vitaladevuni, “Max-pooling loss training of long short-term memory networks for small-footprint keyword spotting,” in 2016 IEEE Spoken Language Technology Workshop (SLT) , 2016, pp. 474–480
2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. C. Loizou, “Speech enhancement: Theory and practice,” 2013
2013
Cited alongside, same era.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” Signal Processing Letters, IEEE , vol. 21, no. 1, pp. 65–68, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
R. Prasad, “Spoken language understanding for amazon echo,” in Keynote in Speech and Audio in the Northeast (SANE) , 2015
2015
Cited alongside, same era.
Later among the works it cites.
S. Panchapagesan, M. Sun, A. Khare, S. Matsoukas, A. Mandal, B. Hoffmeister, and S. Vitaladevuni, “Multi-task learning and weighted cross-entropy for dnn-based keyword spotting.” in Interspeech 2016 , 2016, pp. 760–764
2016
Later among the works it cites.
K. Kumatani, S. Panchapagesan, M. Wu, M. Kim, N. Strom, G. Tiwari, and A. Mandai, “Direct modeling of raw audio with dnns for wake word detection,” in 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2017, pp. 252–257
2017
Later among the works it cites.
J. Guo, K. Kumatani, M. Sun, M. Wu, A. Raju, N. Strom, and A. Mandal, “Time-delayed bottleneck highway networks using a dft feature for keyword spotting,” in ICASSP , 2018
2018
Closest in time.