Fetching the paper…
Reading the bibliography…
In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,”
1989
Earlier work this paper cites.
J. Ortega-García and J. González-Rodríguez, “Overview of speech enhancement techniques for automatic speaker recognition,” in
1996
Earlier work this paper cites.
S. O. Sadjadi and J. H. Hansen, “Assessment of single-channel speech enhancement techniques for speaker identification under mismatched conditions,” in
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-End Factor Analysis for Speaker Verification,”
2011
Earlier work this paper cites.
D. Michelsanti and Z.-H. Tan, “Conditional generative adversarial networks for speech enhancement and noise-robust speaker verification,” in
2012
Earlier work this paper cites.
D. Snyder, G. Chen, and D. Povey, “MUSAN : A Music , Speech , and Noise Corpus,”
2015
Earlier work this paper cites.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-End Text-Dependent Speaker Verification,” in
2016
Earlier work this paper cites.
O. Plchot, L. Burget, H. Aronowitz, and P. Matějka, “Audio enhancing with DNN autoencoder for speaker recognition,” in
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, and Y. Carmiel, “Deep Neural Network Embeddings for Text-Independent Speaker Verification,” in
2017
Cited alongside, same era.
A. Nagraniy, J. S. Chung, and A. Zisserman, “VoxCeleb: A large-scale speaker identification dataset,” in
2017
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep Speaker Recognition,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Shon, H. Tang, and J. Glass, “Frame-level Speaker Embeddings for Text-independent Speaker Recognition and Analysis of End-to-end Model,” in
2018
Later among the works it cites.
H. Tang, W.-N. Hsu, F. Grondin, and J. Glass, “A study of enhancement, augmentation, and autoencoder methods for domain adasptation in distant speech recognition,” in
2018
Later among the works it cites.
W. Cai, J. Chen, and M. Li, “Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,” in
2018
Later among the works it cites.
M. Hajibabaei and D. Dai, “Unified Hypersphere Embedding for Speaker Recognition,”
2018
Later among the works it cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive Statistics Pooling for Deep Speaker Embedding,” in
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D. Bagchi, P. Plantinga, A. Stiff, and E. Fosler-Lussier, “Spectral Feature Mapping With Mimic Loss,” in
2018
Cited alongside, same era.
Later among the works it cites.
J. Villalba, N. Chen, D. Snyder, D. Garcia-Romero, A. McCree, G. Sell, J. Borgstrom, F. Richardson, S. Shon, F. Grondin, R. Dehak, L. P. G. Perera, D. Povey, P. Torres-Carrasquillo, S. Khudanpur, and N. Dehak, “State-of-the-art Speaker Recognition for Telephone and Video Speech: the JHU-MIT Submission for NIST SRE18,” in
2019
Closest in time.
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, “Utterance-level Aggregation For Speaker Recognition In The Wild,” in
2019
Closest in time.