Fetching the paper…
Reading the bibliography…
Research on speaker recognition is extending to address the vulnerability in the wild conditions, among which genre mismatch is perhaps the most challenging, for instance, enrollment with reading speech while testing with conversational or singing audio.
A. E. Rosenberg, “Automatic speaker verification: A review,” Proceedings of the IEEE , vol. 64, no. 4, pp. 475–487, 1976
1976
Earlier work this paper cites.
W. M. Fisher, “Ther DARPA speech recognition research database: specifications and status,” in Proc. DARPA Workshop on Speech Recognition, Feb. 1986 , 1986, pp. 93–99
1986
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on , vol. 1. IEEE Computer Society, 1992, pp. 517–520
1992
Earlier work this paper cites.
T. Matsui and S. Furui, “Concatenated phoneme models for text-variable speaker recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 2. IEEE, 1993, pp. 391–394
1993
Earlier work this paper cites.
D. A. Reynolds, “Automatic speaker recognition using gaussian mixture speaker models,” in The Lincoln Laboratory Journal . Citeseer, 1995
1995
Earlier work this paper cites.
S. Parthasarathy and A. E. Rosenberg, “General phrase speaker verification using sub-word background models and likelihood-ratio scoring,” in Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP’96 , vol. 4. IEEE, 1996, pp. 2403–2406
1996
Earlier work this paper cites.
V. W. Zue and S. Seneff, “Transcription and alignment of the TIMIT database,” in Recent Research Towards Advanced Man-Machine Interface Through Spoken Language . Elsevier, 1996, pp. 515–525
1996
Earlier work this paper cites.
J. P. Campbell, “Speaker recognition: A tutorial,” Proceedings of the IEEE , vol. 85, no. 9, pp. 1437–1462, 1997
1997
Earlier work this paper cites.
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, “Speaker verification using adapted Gaussian mixture models,” Digital signal processing , vol. 10, no. 1-3, pp. 19–41, 2000
2000
Earlier work this paper cites.
D. A. Reynolds, “An overview of automatic speaker recognition technology,” in IEEE international conference on Acoustics, speech, and signal processing (ICASSP) , vol. 4. IEEE, 2002, pp. IV–4072
2002
Earlier work this paper cites.
P. Kenny, M. Mihoubi, and P. Dumouchel, “New MAP estimators for speaker recognition,” in Eighth European Conference on Speech Communication and Technology , 2003
2003
Earlier work this paper cites.
M. P. Alvin and A. Martin, “NIST speaker recognition evaluation chronicles,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop . Citeseer, 2004
2004
Earlier work this paper cites.
P. Kenny, “Joint factor analysis of speaker and session variability: Theory and algorithms,” CRIM, Montreal,(Report) CRIM-06/08-13 , vol. 14, pp. 28–29, 2005
2005
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in European Conference on Computer Vision (ECCV) . Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
N. Brümmer and J. Du Preez, “Application-independent evaluation of speaker detection,” Computer Speech & Language , vol. 20, no. 2-3, pp. 230–275, 2006
2006
Earlier work this paper cites.
P. Kenny, G. Boulianne, P. Ouellet, and P. Dumouchel, “Joint factor analysis versus eigenchannels in speaker recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 4, pp. 1435–1447, 2007
2007
Earlier work this paper cites.
S. J. Prince and J. H. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” in 2007 IEEE 11th International Conference on Computer Vision . IEEE, 2007, pp. 1–8
2007
Earlier work this paper cites.
C. Cieri, L. Corson, D. Graff, and K. Walker, “Resources for new research directions in speaker recognition: The mixer 3, 4 and 5 corpora,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2007
2007
Earlier work this paper cites.
E. Shriberg, M. Graciarena, H. Bratt, A. Kathol, S. S. Kajarekar, H. Jameel, C. Richey, and F. Goodman, “Effects of vocal effort and speaking style on text-independent speaker verification,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2008
2008
Earlier work this paper cites.
D. Ramos and J. Gonzalez-Rodriguez, “Cross-entropy analysis of the information in forensic speaker recognition,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop . International Speech Communication Association, 2008
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
E. Shriberg, S. Kajarekar, and N. Scheffer, “Does session variability compensation in speaker recognition model intrinsic variation under mismatched conditions?” in Tenth Annual Conference of the International Speech Communication Association , 2009
2009
Earlier work this paper cites.
D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y. M. Lui, “Visual object tracking using adaptive correlation filters,” in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . IEEE, 2010, pp. 2544–2550
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. EPFL-CONF-192584. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 4052–4056
2014
Earlier work this paper cites.
J. Gonzalez-Rodriguez, “Evaluating automatic speaker recognition systems: An overview of the nist speaker recognition evaluations (1996-2014),” Loquens , 2014
2014
Cited alongside, same era.
A. Larcher, K. A. Lee, B. Ma, and H. Li, “Text-dependent speaker verification: Classifiers, databases and RSR2015,” Speech Communication , vol. 60, pp. 56–77, 2014
2014
Cited alongside, same era.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,” IEEE Signal processing magazine , vol. 32, no. 6, pp. 74–99, 2015
2015
Cited alongside, same era.
G. Morrison, C. Zhang, E. Enzinger, F. Ochoa, D. Bleach, M. Johnson, B. Folkes, S. De Souza, N. Cummins, and D. Chow, “Forensic database of voice recordings of 500+ australian english speakers,” 2015
2015
Cited alongside, same era.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification.” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2018, pp. 3573–3577
2018
Later among the works it cites.
J. weon Jung, H.-S. Heo, J. ho Kim, H. jin Shim, and H.-J. Yu, “RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 1268–1272
2019
Later among the works it cites.
N. Chen, J. Villalba, and N. Dehak, “Tied mixture of factor analyzers layer to combine frame level representations in neural speaker embeddings,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2948–2952
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. A. Lee, A. Larcher, G. Wang, P. Kenny, N. Brümmer, D. v. Leeuwen, H. Aronowitz, M. Kockmann, C. Vaquero, B. Ma et al. , “The RedDots data collection for speaker recognition,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
L. Li, D. Wang, C. Xing, and T. F. Zheng, “Max-margin metric learning for speaker recognition,” in 10th International Symposium on Chinese Spoken Language Processing (ISCSLP) , 2016, pp. 1–4
2016
Cited alongside, same era.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5115–5119
2016
Cited alongside, same era.
S.-X. Zhang, Z. Chen, Y. Zhao, J. Li, and Y. Gong, “End-to-end attention based text-dependent speaker verification,” in Spoken Language Technology Workshop (SLT) . IEEE, 2016, pp. 171–178
2016
Cited alongside, same era.
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The Speakers in the Wild (SITW) speaker recognition database.” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 818–822
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Cited alongside, same era.
S. J. Park, C. Sigouin, J. Kreiman, P. A. Keating, J. Guo, G. Yeung, F.-Y. Kuo, and A. Alwan, “Speaker identity and voice quality: Modeling human responses and automatic speaker recognition.” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 1044–1048
2016
Cited alongside, same era.
2019
Later among the works it cites.
J. Wang, K.-C. Wang, M. T. Law, F. Rudzicz, and M. Brudno, “Centroid-based deep metric learning for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3652–3656
2019
Later among the works it cites.
Z. Gao, Y. Song, I. McLoughlin, P. Li, Y. Jiang, and L.-R. Dai, “Improving aggregation and loss function for better embedding learning in end-to-end speaker verification system,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 361–365
2019
Later among the works it cites.
J. Zhou, T. Jiang, Z. Li, L. Li, and Q. Hong, “Deep speaker embedding extraction with channel-wise feature responses and additive supervision softmax loss function,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2883–2887
2019
Later among the works it cites.
R. Li, N. Li, D. Tuo, M. Yu, D. Su, and D. Yu, “Boundary discriminative large margin cosine loss for text-independent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6321–6325
2019
Later among the works it cites.
S. Wang, J. Rohdin, L. Burget, O. Plchot, Y. Qian, K. Yu, and J. Cernocky, “On the usage of phonetic information for text-independent speaker embedding extraction,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 1148–1152
2019
Later among the works it cites.
T. Stafylakis, J. Rohdin, O. Plchot, P. Mizera, and L. Burget, “Self-supervised speaker embeddings,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2863–2867
2019
Later among the works it cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4690–4699
2019
Later among the works it cites.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6301–6305
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, and D. Wang, “CN-CELEB: a challenging Chinese speaker recognition dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7604–7608
2020
Closest in time.
Z. Bai, X.-L. Zhang, and J. Chen, “Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6819–6823
2020
Closest in time.
S. Shon and J. Glass, “Multimodal association for speaker verification,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2020, pp. 2247–2251
2020
Closest in time.
J. Kang, R. Liu, L. Li, Y. Cai, D. Wang, and T. F. Zheng, “Domain-invariant speaker vector projection by model-agnostic meta-learning,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2020
2020
Closest in time.
S. Kataria, P. S. Nidadavolu, J. Villalba, and N. Dehak, “Analysis of deep feature loss based enhancement for speaker verification,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2020, pp. 459–466
2020
Closest in time.
J. Mo and L. Xu, “Weighted cluster-range loss and criticality-enhancement loss for speaker recognition,” Applied Sciences , vol. 10, no. 24, p. 9004, 2020
2020
Closest in time.
C. S. Greenberg, L. P. Mason, S. O. Sadjadi, and D. A. Reynolds, “Two decades of speaker recognition evaluation at the national institute of standards and technology,” Computer Speech & Language , vol. 60, p. 101032, 2020
2020
Closest in time.
X. Qin, H. Bu, and M. Li, “Hi-mia: A far-field text-dependent speaker verification database and the baselines,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2020, pp. 7609–7613
2020
Closest in time.
J. W. M. Pham, Z. Li, “Toward better speaker embeddings: Automated collection of speech samples from unknown distinct speakers,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7089–7093
2020
Closest in time.
J. Deng, J. Guo, Y. Zhou, J. Yu, I. Kotsia, and S. Zafeiriou, “Retinaface: Single-stage dense face localisation in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 5203–5212
2020
Closest in time.
D. Wang, “Remarks on optimal scores for speaker recognition,” arXiv preprint arXiv:2010.04862 , 2020
2020
Closest in time.
2020
Closest in time.
Z. Chen, S. Wang, and Y. Qian, “Self-supervised learning based domain adaptation for robust speaker verification,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5834–5838
2021
Closest in time.