Fetching the paper…
Reading the bibliography…
Mismatch between enrollment and test conditions causes serious performance degradation on speaker recognition systems.
C. J. Leggetter and P. C. Woodland, “Maximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models,” Computer speech & language , vol. 9, no. 2, pp. 171–185, 1995
1995
Earlier work this paper cites.
J. P. Campbell, “Speaker recognition: A tutorial,” Proceedings of the IEEE , vol. 85, no. 9, pp. 1437–1462, 1997
1997
Earlier work this paper cites.
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, “Speaker verification using adapted Gaussian mixture models,” Digital signal processing , vol. 10, no. 1-3, pp. 19–41, 2000
2000
Earlier work this paper cites.
D. A. Reynolds, “An overview of automatic speaker recognition technology,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 4. IEEE, 2002, pp. IV–4072
2002
Earlier work this paper cites.
I. J. Myung, “Tutorial on maximum likelihood estimation,” Journal of mathematical Psychology , vol. 47, no. 1, pp. 90–100, 2003
2003
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in European Conference on Computer Vision (ECCV) . Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
N. Brümmer and J. Du Preez, “Application-independent evaluation of speaker detection,” Computer Speech & Language , vol. 20, no. 2-3, pp. 230–275, 2006
2006
Earlier work this paper cites.
P. Kenny, G. Boulianne, P. Ouellet, and P. Dumouchel, “Joint factor analysis versus eigenchannels in speaker recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 4, pp. 1435–1447, 2007
2007
Earlier work this paper cites.
S. J. Prince and J. H. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” in 2007 IEEE 11th International Conference on Computer Vision . IEEE, 2007, pp. 1–8
2007
Earlier work this paper cites.
A. Solomonoff, W. M. Campbell, and C. Quillen, “Nuisance attribute projection,” Speech Communication , pp. 1–73, 2007
2007
Earlier work this paper cites.
E. Shriberg, M. Graciarena, H. Bratt, A. Kathol, S. S. Kajarekar, H. Jameel, C. Richey, and F. Goodman, “Effects of vocal effort and speaking style on text-independent speaker verification,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2008
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
N. Brümmer and E. De Villiers, “The speaker partitioning problem.” in Odyssey , 2010, p. 34
2010
Earlier work this paper cites.
L. Wang and T. F. Zheng, “Creation of time-varying voiceprint database,” in Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 2010, pp. 1–5
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
M. McLaren and D. Van Leeuwen, “Source-normalized lda for robust speaker recognition using i-vectors from multiple speech sources,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 3, pp. 755–766, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. EPFL-CONF-192584. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
J. Villalba and E. Lleida, “Bayesian adaptation of PLDA based speaker recognition to domains with scarce development data,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2012, pp. 47–54
2012
Earlier work this paper cites.
Y. Lei, L. Burget, L. Ferrer, M. Graciarena, and N. Scheffer, “Towards noise-robust speaker recognition using probabilistic linear discriminant analysis,” in 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2012, pp. 4253–4256
2012
Earlier work this paper cites.
M. I. Mandasari, R. Saeidi, M. McLaren, and D. A. van Leeuwen, “Quality measure functions for calibration of speaker recognition systems in various duration conditions,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 21, no. 11, pp. 2425–2438, 2013
2013
Earlier work this paper cites.
P. Kenny, T. Stafylakis, P. Ouellet, M. J. Alam, and P. Dumouchel, “PLDA for speaker verification with utterances of arbitrary duration,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 7649–7653
2013
Earlier work this paper cites.
B. J. Borgström and A. McCree, “Discriminatively trained bayesian speaker comparison of i-vectors,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 7659–7662
2013
Earlier work this paper cites.
P. Rajan, T. Kinnunen, and V. Hautamäki, “Effect of multicondition training on i-vector plda configurations for speaker recognition.” in Interspeech . Citeseer, 2013, pp. 3694–3697
2013
Earlier work this paper cites.
L. Deng and D. Yu, “Deep learning: methods and applications,” Foundations and trends in signal processing , vol. 7, no. 3–4, pp. 197–387, 2014
2014
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 4052–4056
2014
Earlier work this paper cites.
D. Garcia-Romero and A. McCree, “Supervised domain adaptation for i-vector based speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 4047–4051
2014
Earlier work this paper cites.
D. Garcia-Romero, A. McCree, S. Shum, and C. Vaquero, “Unsupervised domain adaptation for i-vector speaker recognition,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2014
2014
Cited alongside, same era.
O. Glembek, J. Ma, P. Matějka, B. Zhang, O. Plchot, L. Bürget, and S. Matsoukas, “Domain adaptation via within-class covariance correction in i-vector based speaker recognition systems,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 4032–4036
2014
Cited alongside, same era.
H. Aronowitz, “Inter dataset variability compensation for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 4002–4006
2014
Cited alongside, same era.
H. Aronowitz, “Compensating inter-dataset variability in PLDA hyper-parameters for robust speaker recognition.” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2014, pp. 280–286
2014
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2018, pp. 1086–1090
2018
Later among the works it cites.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2018, pp. 74–81
2018
Later among the works it cites.
W. Ding and L. He, “MTGAN: Speaker verification through multitasking triplet generative adversarial networks,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2018, pp. 3633–3637
2018
Later among the works it cites.
Q. Wang, W. Rao, S. Sun, L. Xie, E. S. Chng, and H. Li, “Unsupervised domain adaptation via domain adversarial training for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4889–4893
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. McLaren, A. Lawson, L. Ferrer, N. Scheffer, and Y. Lei, “Trial-based calibration for speaker recognition in unseen conditions,” in Proc. Odyssey , 2014, pp. 19–25
2014
Cited alongside, same era.
A. Sizov, K. A. Lee, and T. Kinnunen, “Unifying probabilistic linear discriminant analysis variants in biometric authentication,” in Joint IAPR International Workshops on Statistical Techniques in Pattern Recognition (SPR) and Structural and Syntactic Pattern Recognition (SSPR) . Springer, 2014, pp. 464–475
2014
Cited alongside, same era.
2014
Cited alongside, same era.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,” IEEE Signal processing magazine , vol. 32, no. 6, pp. 74–99, 2015
2015
Cited alongside, same era.
A. Kanagasundaram, D. Dean, and S. Sridharan, “Improving out-domain PLDA speaker verification using unsupervised inter-dataset variability compensation approach,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4654–4658
2015
Cited alongside, same era.
2015
Cited alongside, same era.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5115–5119
2016
Cited alongside, same era.
S.-X. Zhang, Z. Chen, Y. Zhao, J. Li, and Y. Gong, “End-to-end attention based text-dependent speaker verification,” in Spoken Language Technology Workshop (SLT) . IEEE, 2016, pp. 171–178
2016
Cited alongside, same era.
2018
Later among the works it cites.
M. H. Rahman, A. Kanagasundaram, I. Himawan, D. Dean, and S. Sridharan, “Improving PLDA speaker verification performance using domain mismatch compensation techniques,” Computer Speech & Language , vol. 47, pp. 240–258, 2018
2018
Later among the works it cites.
C. Zhang, S. Ranjan, and J. H. Hansen, “An analysis of transfer learning for domain mismatched text-independent speaker verification,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2018, pp. 181–186
2018
Later among the works it cites.
L. Ferrer and M. McLaren, “A generalization of plda for joint modeling of speaker identity and multiple nuisance conditions.” in INTERSPEECH , 2018, pp. 82–86
2018
Later among the works it cites.
J. weon Jung, H.-S. Heo, J. ho Kim, H. jin Shim, and H.-J. Yu, “RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 1268–1272
2019
Later among the works it cites.
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, “Utterance-level aggregation for speaker recognition in the wild,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5791–579
2019
Later among the works it cites.
N. Chen, J. Villalba, and N. Dehak, “Tied mixture of factor analyzers layer to combine frame level representations in neural speaker embeddings,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2948–2952
2019
Later among the works it cites.
J. Wang, K.-C. Wang, M. T. Law, F. Rudzicz, and M. Brudno, “Centroid-based deep metric learning for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3652–3656
2019
Later among the works it cites.
Z. Gao, Y. Song, I. McLoughlin, P. Li, Y. Jiang, and L.-R. Dai, “Improving aggregation and loss function for better embedding learning in end-to-end speaker verification system,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 361–365
2019
Later among the works it cites.
J. Zhou, T. Jiang, Z. Li, L. Li, and Q. Hong, “Deep speaker embedding extraction with channel-wise feature responses and additive supervision softmax loss function,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2883–2887
2019
Later among the works it cites.
R. Li, N. Li, D. Tuo, M. Yu, D. Su, and D. Yu, “Boundary discriminative large margin cosine loss for text-independent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6321–6325
2019
Later among the works it cites.
S. Wang, J. Rohdin, L. Burget, O. Plchot, Y. Qian, K. Yu, and J. Cernocky, “On the usage of phonetic information for text-independent speaker embedding extraction,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 1148–1152
2019
Later among the works it cites.
T. Stafylakis, J. Rohdin, O. Plchot, P. Mizera, and L. Burget, “Self-supervised speaker embeddings,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 2863–2867
2019
Later among the works it cites.
L. Ferrer and M. McLaren, “Joint plda for simultaneous modeling of two factors,” The Journal of Machine Learning Research , vol. 20, no. 1, pp. 847–875, 2019
2019
Later among the works it cites.
J. Williams and S. King, “Disentangling style factors from speaker representations,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2019, pp. 3945–3949
2019
Later among the works it cites.
G. Bhattacharya, J. Alam, and P. Kenny, “Adapting end-to-end neural speaker verification to new languages and recording conditions with adversarial training,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6041–6045
2019
Later among the works it cites.
K. A. Lee, Q. Wang, and T. Koshinaka, “The CORAL+ algorithm for unsupervised domain adaptation of plda,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5821–5825
2019
Later among the works it cites.
X. Qin, H. Bu, and M. Li, “HI-MIA: A far-field text-dependent speaker verification database and the baselines,” 2019
2019
Later among the works it cites.
Z. Bai, X.-L. Zhang, and J. Chen, “Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6819–6823
2020
Closest in time.
2020
Closest in time.
D. Wang, “Remarks on optimal scores for speaker recognition,” CSLT@Tsinghua University, 2020. [Online]. Available: http://wangd.cslt.org/public/pdf/nl.pdf
2020
Closest in time.
D. Wang, “A simulation study on optimal scores for speaker recognition,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2020, no. 1, pp. 1–23, 2020
2020
Closest in time.
W. H. Kang, S. H. Mun, M. H. Han, and N. S. Kim, “Disentangled speaker and nuisance attribute embedding for robust speaker verification,” IEEE Access , 2020
2020
Closest in time.