Fetching the paper…
Reading the bibliography…
Automatic speech recognition (ASR) is a key technology in many services and applications.
D. A. Reynolds, “Speaker identification and verification using Gaussian mixture speaker models,”
1995
Earlier work this paper cites.
K. Sekiyama, “Cultural and linguistic factors in audiovisual speech processing: The McGurk effect in Chinese subjects,”
1997
Earlier work this paper cites.
A. Stolcke, E. Shriberg, R. Bates, N. Coccaro, D. Jurafsky, R. Martin, M. Meteer, K. Ries, P. Taylor, C. Van Ess-Dykema
1998
Earlier work this paper cites.
A. A. Dibazar, S. Narayanan, and T. W. Berger, “Feature analysis for automatic detection of pathological speech,” in
2002
Earlier work this paper cites.
O.-W. Kwon, K. Chan, J. Hao, and T.-W. Lee, “Emotion recognition by speech signals,” in
2003
Earlier work this paper cites.
D. Ververidis and C. Kotropoulos, “Automatic speech classification to five emotional states based on gender information,” in
2004
Earlier work this paper cites.
K. Umapathy and S. Krishnan, “Feature analysis of pathological speech signals using local discriminant bases technique,”
2005
Earlier work this paper cites.
Y.-M. Zeng, Z.-Y. Wu, T. Falk, and W.-Y. Chan, “Robust GMM based gender classification using pitch and RASTA-PLP parameters of speech,” in
2006
Earlier work this paper cites.
M. Kotti and C. Kotropoulos, “Gender classification in two emotional speech databases,” in
2008
Earlier work this paper cites.
A. Vinciarelli, M. Pantic, and H. Bourlard, “Social signal processing: Survey of an emerging domain,”
2009
Earlier work this paper cites.
M. El Ayadi, M. S. Kamel, and F. Karray, “Survey on speech emotion recognition: Features, classification schemes, and databases,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
M. A. Pathak,
2012
Earlier work this paper cites.
T. Ballmer and W. Brennstuhl,
2013
Cited alongside, same era.
B. Schuller, S. Steidl, A. Batliner, A. Vinciarelli, K. Scherer, F. Ringeval, M. Chetouani, F. Weninger, F. Eyben, E. Marchi
2013
Cited alongside, same era.
B. Schuller and A. Batliner,
2013
Cited alongside, same era.
B. Schuller, S. Steidl, A. Batliner, E. Nöth, A. Vinciarelli, F. Burkhardt, R. van Son, F. Weninger, F. Eyben, T. Bocklet
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
D. Snyder, G. Chen, and D. Povey, “MUSAN: A music, speech, and noise corpus,”
Y. Gu, X. Li, S. Chen, J. Zhang, and I. Marsic, “Speech intention classification with multimodal deep learning,” in
2017
Later among the works it cites.
C. Glackin, G. Chollet, N. Dugan, N. Cannings, J. Wall, S. Tahir, I. G. Ray, and M. Rajarajan, “Privacy preserving encrypted phonetic search of speech data,” in
2017
Later among the works it cites.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,”
2017
Later among the works it cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in
2017
Later among the works it cites.
V. Kepuska and G. Bohouta, “Next-generation of virtual personal assistants (Microsoft Cortana, Apple Siri, Amazon Alexa and Google Home),” in
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Cited alongside, same era.
N. Hellbernd and D. Sammler, “Prosody conveys speaker’s intentions: Acoustic cues for speech act perception,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,”
2016
Cited alongside, same era.
G. López, L. Quesada, and L. A. Guerrero, “Alexa vs. Siri vs. Cortana vs. Google Assistant: a comparison of speech-based natural user interfaces,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in
2018
Later among the works it cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Tsuchiya, N. Tawara, T. Ogawa, and T. Kobayashi, “Speaker invariant feature extraction for zero-resource languages with adversarial learning,” in
2018
Later among the works it cites.
Z. Meng, J. Li, Z. Chen, Y. Zhao, V. Mazalov, Y. Gong
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.